When it comes to smooth e-commerce capability in the modern era, Commerce provides the technology behind the technology—the platform that’s ultimately the backbone of buying for tens of thousands of merchants, including Solo Stove, Skullcandy, Yeti, and a lot of other incredibly popular brands and products.
For a company that enables varying levels of demand for merchants in the digital age, timeliness and accuracy of data are crucial. Inventory has to be accurate to the millisecond, and when data is delayed by a day, or even a minute, it results in lost revenue for merchants and damages to the organization’s reputation.
That’s only one of the reasons Commerce pursued Apache Kafka for their data streaming use case. The modern, open SaaS, e-commerce platform works with tens of thousands of merchants who serve millions of customers around the world. Its digital commerce solutions are available 24/7, all year long, for B2B, B2C, multi-storefront, international, and omnichannel customers.
Recognizing the need for real-time data while understanding the burden of self-managing Kafka on their own led Commerce to choose Confluent—allowing them to tap into data streaming without having to manage and maintain the data infrastructure.
The challenges of self-managed Kafka when instant access to data insights is critical to business
Before Confluent, Commerce was managing its own Kafka cluster that was growing and required increasing maintenance. Commerce started out with six broker nodes, but as their traffic and use cases increased, they ended up trying to manage 22 broker nodes and five ZooKeeper nodes. This required over 20 hours a week dedicated to maintaining and managing Kafka infrastructure and about half a full-time engineer’s time.
They chose Kafka because it is a robust, resilient system designed for high throughput, which would allow the company’s data teams to retain data for a period of time, as well as offset and reprocess data as needed. However, open source (OSS) Kafka presented significant limitations. Rather than being able to focus on delivering new services that improve user experiences, team members were bogged down by software patches, blind spots in data-related infrastructure, and system updates.
Challenge #1: Maintaining and managing Kafka
Having the latest and greatest improvements to Kafka is important, and Commerce wanted to get the most out of the infrastructure that was running Kafka. However, managing upgrade priorities had to be balanced with other roadmap features that the team needed to deliver.
Once the company got up to 22 Kafka nodes, management became unwieldy. The engineering team was composed of nine people to manage the data platform, data applications, data APIs, analytics, and data infrastructure. However, they had three engineers spending half their time on Kafka-related maintenance and upgrades.
Before big shopping events like CyberWeek, the entire team needed a month to get the technology ready to scale and had to overprovision their clusters by 10-15% to account for the uncertain traffic volume. This exercise was expensive and time-consuming, and valuable resources were being squandered because of the unknowns.
Challenge #2: Inability to tap into real-time data creates burden for customers
Kafka was originally instituted to help Commerce capture critical retail events such as orders and cart updates. These events were stored in a NoSQL database as they came in. But with that model, the company was not able to tap into streaming data, and that meant they couldn’t exactly provide real-time data analysis for merchants who wanted to be able to extract quick insights from the data.
From time to time, the system would go down, and that would cause a spike of events to backlog, which degraded the merchant and end-user experience. Commerce had an ETL, batch-based system for analytics and insights, resulting in a wait time of eight hours before merchants could get their analytics reports. This system consisted of 30+ MapReduce jobs. On occasion, jobs would fail and need manual intervention, causing even further delays.
The decision to adopt a fully managed Kafka solution
At the time of migration, Commerce processed around 1.6B events a day. These were high-traffic, high-value events such as order events, cart events, and page view events. The team knew that with Kafka, they already had the ability to consume data in real time. They just needed the support of a fully managed data streaming platform to help them tap into the potential of Kafka.
In addition, the company’s storefront architecture was already on Google Cloud, so the fact that Confluent Cloud integrates so easily with GCP was a key lever. Using Confluent on Google Cloud allowed Commerce to significantly cut down on data transfer costs.
Last but not least, Commerce needed the platform to enable them to stream data to their own customers—the merchants—as part of their valuable open SaaS strategy. This would enable merchants to integrate with third-party systems such as Google Analytics and other APIs to, for instance, run ad campaigns for custom audiences.
Migrating from self-managed Kafka to Confluent Cloud
When Commerce wanted to migrate from their existing open source Kafka to Confluent Cloud, they embarked upon a phased approach to ensure a smooth transition without any impact to their critical workloads.
The team meticulously planned the migration with the help of Confluent’s expertise to design the partition strategy and migration solution. This was done to prepare for future needs and traffic variations for production use cases.
The first phase of the migration saw the piped production data being sent to both OSS Kafka and Confluent to monitor latency and avoid any possible downtime. Once the systems were fully in sync, the switch to Confluent was seamless.
In the second phase, Commerce integrated consumer lag and sink connector lag reporting into a single unified Google Cloud monitoring dashboard with the help of Confluent Metrics API. This helped them keep track of the lagging issues and provided a way to address them in a timely manner.
The third phase involved the migration of insights and data lake workloads, loading data into Google BigQuery using Confluent Sink Connector. Finally, in the fourth phase, Commerce optimized Kafka stream performance and determined the optimal number of consumers for future workloads.
In five months, Commerce had migrated 1.6 billion events per day, 22 broker nodes, three clusters, and 15 TB of storage data to Confluent Cloud. They experienced zero downtime, zero data loss, and zero disruption to the merchant experience.
With Confluent Cloud and fully managed Kafka, attention turns to innovation
The resulting benefits of migration for Commerce’s engineering team have had a strong impact on the business, the merchant, and the end-user experience.
The biggest benefit is the simplification of using Kafka. With Kafka fully managed in Confluent Cloud, the engineering team no longer has to spend 20 hours a week managing low-level Kafka infrastructure—upgrades, traffic spikes, compliance, security, nodes, and other tasks. They also haven’t had to overprovision clusters by 10-15% “just in case,” because Confluent is elastically scalable. Instead, they have been able to turn their full attention to innovating products in a highly competitive industry.
In addition, Confluent enables the team to design products and features around asynchronous communication. Some systems consume immediately, and some consume at other times.
For merchants, the switch to Confluent means they can get real-time analytics on Commerce’s open platform and tap into precise insights for their specific use cases. The data coming out of Commerce is now trustworthy, timely, accurate, and actionable. Because Confluent is cloud native and has the ability to “connect everywhere” via integrations, Commerce is able to take real advantage of the data pipeline.
They’re now pulling in data via the flexibility and ease of Confluent’s Google Cloud Storage Sink Connector, and merchants can easily combine this real-time data with the data they already have in their ecosystems so they can get to the ideal state of a 360-degree view.
The challenges of building a real-time solution on Kafka are not always easy to solve, particularly for an in-house engineering team that’s better off spending time on product innovation versus techstack management. Kafka self-management requires people, time, and effort—and can be a process of trial and error.
Business Results
Reduced operational burden - The team no longer worries about managing Kafka and instead focuses on feature functionality. Specifically, there’s no longer a need for the data team to spend 20 hours a week managing low level Kafka infrastructure.
Faster time to market - As artificial intelligence (AI) and other technologies become competitive levers for retail companies, Commerce can now get things like a product-recommendation proof of concept built faster.
The ability to design for asynchronicity - Confluent enables Commerce to build products and features that enable every system to consume messages in real time or not.
Cost savings - The simple fact that Commerce no longer has to assign engineers to manage Kafka creates cost savings, specifically around the operational overhead that had previously been required to scale, patch, upgrade, monitor, rebalance, and optimize the platform.
Technical Results
Elastic scalability - Confluent has the ability for high throughput, so it handles traffic seamlessly, even during seasonal superspikes like CyberWeek. There’s no need to overprovision clusters by 10-15% “just in case” like they used to.
No downtime - Since completing the migration to Confluent Cloud, there has been zero data loss and zero downtime.
Integration with third-party systems - The ability to connect streaming data from Kafka with Confluent’s many connectors provides great engineering flexibility and opportunity.
Jetzt mit Confluent loslegen
Jetzt bei neuer Registrierung Credits im Wert von 400 $ erhalten, die in den ersten 30 Tagen genutzt werden können.



