Big Data Consulting to Speed Up Reporting for a US Retailer

Big Data Consulting to Speed Up Reporting for a US Retailer

Industry
Retail
Technologies
MS SQL Server, Python, Spark, Big data, Hadoop
Business gains
100x faster big data processing

Summary

INNERLUXES performed an audit of an enterprise-wide big data reporting solution for a US apparel manufacturer and provided recommendations that helped significantly boost the system's performance.

About the Client

The Client is a US-based manufacturer and distributor of branded and private-label clothing for children and adults, with more than 4,000 employees worldwide.

The Client had an on-premises reporting solution based on Microsoft SQL Server that it relied on for daily sales reports. As the business grew, the solution could not accommodate the increasing data volume, so the company upgraded the system with HDFS, Apache Spark, and Apache Hive technologies. However, the data processing speed remained subpar, which caused delays in the delivery of business-critical insights.

The Client needed the solution to process more than 300 GB of ORC files every day to provide up-to-date sales insights, and the compound dataset included billions of historical records that were expanded with daily updates. The company expected the number of records to double within the next two years.

To address the pressing data processing issues and find a future-proof solution to the challenge, the Client was looking for an IT vendor with big data experience to review the reporting system and offer actionable improvement recommendations.

Big Data Solution Audit and Consultation Sessions

INNERLUXES held several interviews with the Client's stakeholders to understand the company's reporting needs, learn about the configurations of the big data solution, and assess the attempted remediation measures.

INNERLUXES concluded that the Client's system needed a revamp to enable timely reporting and accommodate future data volume growth. However, given the business-critical nature of the reporting solution, the team suggested optimizing the existing system first and carrying out a full revamp in the future.

For six weeks, INNERLUXES held sessions with five IT specialists from the Client's team — advising on the optimal remediation steps, gathering feedback on the achieved results, and providing instructions on further relevant measures, along with multiple Q&A and knowledge-sharing sessions.

Acting on INNERLUXES's advice, the Client achieved the following improvements:

  • Implemented efficient partitioning and partition pruning of database tables, which increased the execution speed of BI queries performed using Hive.
  • Replaced MapReduce with a properly configured Tez execution engine for Hive processing, ensuring more efficient allocation of memory and CPU resources across tasks.
  • Moved part of the logic to the Spark SQL cluster to enable ACID transformations.
  • Fine-tuned Hive and Spark configurations.
  • Used MSC repair, Spark compact files, and bucketing to optimize data storage for high-performance read operations.

The Client also received a detailed report on the pros and cons of further improving the current on-premises solution versus building a new cloud-based system.

Key Outcomes

  • A complete audit of the business-critical reporting solution completed within six weeks.
  • 100x faster big data processing for timely sales insights and forecasts.
  • The foundation for a gradual solution revamp without business disruptions. The Client now uses the improved reporting system and continues to collaborate with INNERLUXES on developing a cloud-based big data analytics solution.

Technologies and Tools

HDFS, Apache Spark, Apache Hive, PySpark, Python, Microsoft SQL Server, T-SQL.