Custom RPA Development for ECM System Efficiency

Custom RPA Development for ECM System Efficiency

Industry
AI/ML, Intelligent Document Capture
Core Technologies
Regex, Artificial Intelligence, Machine Learning

Summary

INNERLUXES delivered automated document classification and data extraction using RPA, OCR, and Ephesoft to streamline enterprise content management (ECM) integration for a provider of intelligent document capture technology.

Robotic Process Automation Software

Just as digitalization has reached ubiquity, automation is following a similar path through Robotic Process Automation (RPA) and Artificial Intelligence (AI) technologies. By defining a context for a series of inputs, automation technologies dictate the corresponding outputs, produced automatically from predefined parameters. The most obvious difference between AI and RPA is their ability to evolve: AI can evolve past the original parameters to produce new outputs beyond the rules explicitly dictated, while RPA produces rule-bound outputs.

RPA improves business processes by eliminating human error, and by removing human capital from repetitive tasks it reduces operating costs. The Client uses RPA technology to reduce both resource and process inefficiencies. Its intelligent document capture technology classifies and separates page images into documents and extracts metadata from the Optical Character Recognition (OCR) content of a document.

The Challenge

The Client's web-based user interface allows operators to review documents and validate the extracted content. The assembled document and its associated metadata can then be exported to other ECM systems for further processing. All organizations require tools to manage their information and knowledge — document management, workflow, web content management, document capture, records management, and portals are a few of the tools categorized as ECM. Because information can be stored in many different electronic systems, ECM tools must communicate not only with each other but also with corporate systems such as ERP, CRM, and other associated databases.

The Solution

INNERLUXES's RPA developer team engineered a document capture and data classification platform capable of accelerating the identification of various documents and extracting meaningful data to feed back-office applications and business processes. Ephesoft software is used to extract the desired data.

Within the data extraction process, the supporting software must be trained with extraction rules that create the operating requirements. These rules are set as text patterns — matching regular expressions to identify small- or large-scale structures. After extraction, each rule is validated, and once the desired result is obtained the rule can be applied; each subsequent extraction then mirrors the rule to produce the desired result.

Extraction Rules Secure Accurate RPA Outputs

The RPA developer builds the environment on Ephesoft software. The rules (or codes) are based on the classification of the documents. Batch classes are created after classification, and each document is indexed where extraction rules are established. There are three types of extraction:

  • Free Form Extraction
  • Fixed Form Extraction
  • Table / Line Item Extraction

The RPA developer uses Regular Expressions (Regex) in the rule set-up, and Ephesoft software supports the dictated rules to jump-start the desired extraction. An administrator acts as the Quality Analyst and validates the extracted data to ensure it aligns with the customer's requirements. If the extraction is not cohesive, the administrator re-assigns the work to a Data Extraction team, which completes the process manually.

For data extraction approved by the administrator, the corresponding output is imported into the immediate extracting system. This is followed by data transformation and, where needed, the addition of metadata before exporting to the next stage in the workflow. Multiple users can access the extraction and validation. Once sample files are provided to create extraction rules in one format, the rules can be applied in an automated extraction. Each desired extraction must be written into a rule requirement by an RPA developer to operate correctly.

The Client lists all the document types and fields to be extracted; the RPA developer then builds enumerations into the system so that extractions are obtained with greater accuracy. The Regex library accommodates all possible fields, and document types are described by annotations within the document.

The Results

  • The Ephesoft software becomes more attuned to the Client's needs over time — the more it is trained, the more accurate it becomes.
  • The Client improved nearly every facet of its infrastructure, rendering repetitive manual processes obsolete.
  • The engagement evolved from a single project into a long-term collaboration.

Technologies and Tools

Ephesoft, OCR, Regex, Robotic Process Automation (RPA), Artificial Intelligence, Machine Learning.