Cradl AI Transforms Unstructured Documents into Structured Data with Advanced AI Technology
In today's data-driven landscape, organizations generate an overwhelming volume of unstructured documents that contain critical information. From complex financial invoices to legal agreements, extracting actionable data from these documents remains a time-consuming challenge. Enter Cradl AI, a powerful document automation platform that leverages advanced AI technology to transform raw documents into structured data. In this technical exploration, we'll uncover how Cradl processes everything from simple PDFs to intricate multi-page invoices, extracting key information with remarkable accuracy. Through its custom machine learning architecture and sophisticated validation platform, the system achieves unprecedented automation while maintaining high standards of data quality. Join us as we dive into the technical details of Cradl's document processing capabilities, including its integrations with popular automation tools and its robust security framework.
Through its AI-driven workflows, Cradl processes unstructured documents into actionable data, handling everything from PDFs to complex invoices with diverse layouts. The platform supports multiple document creation methods, including direct upload, API integration, and email data extraction, making it versatile for various implementation scenarios.
The engine that powers Cradl's capabilities is its custom machine learning architecture, which learns from user-specific data without requiring rigid templates or structured layouts. This flexibility allows it to extract key information from a wide range of documents, with each model trained to understand specific document types and languages.
At the core of Cradl's functionality is its advanced validation platform, which assigns certainty scores to each extracted piece of information. Ranging from 0 to 100, these scores help users balance automation and accuracy by allowing customizable thresholds for manual review. This feature, combined with tools for ongoing model training and improvement, represents a significant advancement in document automation technology.
Cradl seamlessly integrates with popular automation tools like Power Automate, Zapier, and email systems to extract data directly from document attachments. The platform supports complex document structures, including tables and line items, with robust handling capabilities for files exceeding 70 pages. This strength in processing complex documents sets it apart from many competing AI tools in the market.
Users create custom machine learning models through structured data bundles, specifying parameters like image resolution and field configurations. The platform continuously learns from user data, maintaining high accuracy without the need for constant intervention. Each model undergoes rigorous quality assessment before deployment, with detailed error codes and reporting tools helping users understand and improve their data quality.
Cradl's AI models learn from unstructured documents without requiring predetermined templates or layouts. The platform supports multiple input methods, including direct file uploads, API integrations, email attachment extraction, and document management system (DMS) connections (Power Automate, Zapier, email systems). Each document type undergoes AI processing to extract key information, with the system achieving over 90% reduction in processing time for complex documents containing tables or line items (up to 70 pages).
The extraction process assigns a certainty score between 0 and 100 to each piece of extracted information, allowing users to set custom thresholds for manual review (industry average is 70-100%). This feature enables users to balance automation and accuracy based on their specific requirements. The system provides detailed error codes for misidentifications, including "DE001" for generic document processing errors and "LE001" for generic label parsing errors, facilitating quick error resolution.
Document processing requires specific parameters defined in JSON format, including image width, height, and field configuration. Training data must meet several quality criteria: sufficient quantity (15 documents per label with 15 ground truths per label), good variation, and representativeness across document types and languages. Recommended image sizes are multiples of 320 plus one pixel, with optimal performance for A4 documents with fine print (1281x961) and square documents with larger print (321x321).
The platform's security framework employs Amazon Web Services (AWS) best practices, including TLS encryption for data transport and AES-256 encryption at rest in GCM mode. All encryption keys are managed by AWS Key Management Service (KMS) with unique keys per account. Data storage duration follows user-defined policies up to 12 months, with automated lifecycle management rules and timers. AWS Certificate Manager handles all certificates, and the platform uses Amazon API Gateway's DoS-protection mechanism along with various throttling limits to prevent abuse. Regular automated health checks ensure service availability across multiple availability zones.
Before training a model, users must create a Data Bundle, which acts as a model-specific collection of one or more Datasets. The process begins by specifying the modelId and one or more datasetIds, with options to add a name and description. Each dataset consists of one or more documents with associated ground truth information.
The system automatically generates a Data Report upon Data Bundle creation, evaluating the quality of the included data across six statistical measures: completeness, validity, coverage, uniqueness, uniformity, and variation. These metrics collectively determine the overall data quality score, with an ideal range of 70-100%. The report generation process may take several minutes, particularly for larger datasets.
To successfully train a model, the data must meet three key criteria: sufficient quantity (at least 15 documents per label with 15 ground truths per label), good variation across labels, and representativeness covering all relevant document types and languages. The model's performance is significantly impacted by the variability of training examples. For instance, if a majority of additional invoices come from a single company, the model may learn to prioritize predictions for that company's name, even if such examples represent 90% of the training data.
The system's training process involves both vision and textual understanding, handling documents with complex layouts and structures. Supported languages include those within the extended Latin alphabet, with optimal performance demonstrated for A4 documents with fine print (1281x961) and square documents with larger print (321x321). Image quality is specified in JSON format through the preprocessConfig parameter, which controls aspects like auto-rotation and maximum page count.
Each model's performance can be customized through field configurations defined in a JSON file, specifying extraction details including field names and data formats. Supported data formats include dates (YYYY-MM-DD), amounts (with two decimal places), strings, digits, and enums. Training data must cover all relevant document types and languages, with demonstrated requirements for both text and numerical inputs.
Cradl supports document processing through multiple creation methods, including Command Line Interface (CLI), cURL, Python SDK, and direct file uploads. Each document consists of image or PDF content along with associated ground truth information, which can be specified during creation or updated later. Ground truth data is provided as a JSON array of objects containing label and value pairs, with an emphasis on maintaining label consistency across documents and models.
The platform offers comprehensive document processing capabilities, supporting various formats including PDF, JPEG, PNG, and TIFF. It integrates with popular automation tools through native integrations with Power Automate and Zapier, while also providing email extraction capabilities directly from attachments. For data management, users have full control over retention policies, with data stored for up to 12 months by default. The company uses Amazon Web Services (AWS) to manage data transport and storage security, implementing AWS best practices for encryption and security compliance.
Cradl provides detailed error codes and reporting mechanisms to help users understand and rectify issues, including "DE001" for generic document processing errors and "LE001" for generic label parsing errors. The platform continuously monitors system health across multiple availability zones and implements automated lifecycle management rules to ensure data retention complies with user-defined policies. To protect user data, Cradl does not store credit card information and uses Stripe as a payment processor, maintaining all data within customer control.
The company offers tiered pricing plans to accommodate different organization sizes and needs, with individual plans starting at $40 per month for 200 pages. For larger organizations, the professional plan provides 1000 pages at $0.25 per page and advanced features like priority support and custom security measures. Enterprise customers have access to custom pricing models and dedicated customer success managers. All plans include support for multiple automation integrations and no-base model training, with fine-tuned AI models available at no additional cost.
Cradl employs a comprehensive security framework centered around Amazon Web Services (AWS), implementing best practices for data transport and storage security. All data transport, including communication between virtual private clouds (VPCs), utilizes Transport Layer Security (TLS) encryption. Data at rest is protected using AES-256 encryption in Galois/Counter Mode (GCM), with all encryption keys managed by AWS Key Management Service (KMS) on a per-account basis.
The platform's security architecture emphasizes immutable infrastructure principles, with developers lacking direct access to production environments, including S3 storage buckets. Any modifications to production code require version control, review, and testing via a pipeline that combines CodePipeline, CodeBuild, and CloudFormation. Cradl adheres to least privilege principles in all infrastructure access controls.
Physical security is managed by AWS infrastructure, which provides redundancy, scalability, and key management capabilities. The company conducts annual security testing using a suite of tools including AWS CloudTrail, AWS WAF, AWS GuardDuty, and Snyk to monitor and detect threats and vulnerabilities. Incident management protocols are in place, though the company acknowledges the potential for residual vulnerabilities due to continuous operations.
All employees are physically located in Europe, particularly Norway, and maintain encrypted hard drives on their work laptops. For trips to high-risk countries, they use clean devices to prevent potential security breaches. The company employs Amazon Certificate Manager (ACM) for certificate management and AWS API Gateway's built-in Distributed Denial of Service (DDoS) protection mechanisms alongside various throttling limits to prevent service abuse.
User data management follows detailed retention policies, with storage duration determined by user-defined policies that can extend up to 12 months. This period serves dual purposes: document processing and model training improvement. Data lifecycle management employs AWS's automated rules, hooks, and timers to enforce these policies. Users retain full control over document deletion through Cradl's APIs, with the option to use a hashed end-user ID as a consent identifier for third-party processing.
The platform processes documents on behalf of multiple clients, implementing strict sub-processor controls. AWS serves as the sole sub-processor for all personal data processing, which occurs entirely within the European Union. All third-party service providers, including those for analytics and transactional emails, are disclosed in the platform's documentation. Security testing is conducted regularly, with the company maintaining a service-level agreement (SLA) to govern performance commitments.