All resources

The Data Provider's Guide to Modernizing Delivery

Delivery is now a differentiator for data businesses. Learn why and how leading data providers are modernizing their delivery programs beyond SFTP.

Bobsled Team

Bobsled Team

April 23, 2023

Delivery is now a differentiator for data businesses. After decades of using SFTP and APIs, customers are demanding that providers share data directly into the cloud storage, data warehouses and other cloud platforms where they work. This guide provides data providers with an introduction to cloud sharing and offers a set of best practices to modernize your delivery program — whether you are just exploring your first cloud destination or investing in scaling your program.

As data operations scale, the cost and pain of integrating external data through traditional technologies like SFTP or APIs is becoming a strategic barrier for buyers and an anchor on the data industry as a whole. Cloud sharing — in which a provider uses the native sharing protocols developed by the major platforms to deliver analytics-ready data directly into the cloud data lake or warehouse where their customers operate — is quickly becoming a favorite delivery mechanism for modern data teams and providers alike. A new generation of data sharing platforms is eliminating much of the pain of scaling a cloud delivery program.

Why modernizing delivery is essential to scaling data businesses

The external data industry is at an important crossroads. On one hand, demand for data has never been stronger, as data teams — empowered by their role as a strategic asset — invest in products that drive value for their organizations. On the other hand, as data teams have scaled, the way they consume products is changing, and delivery is now a critical consideration for many buyers.

The search for efficiency is at the heart of these new requirements. Manual, time-consuming processes such as integration — which for years were accepted by data consumers — are quickly becoming major barriers to growth for both buyers and sellers. Buyers now demand that providers find more scalable ways to discover and deliver their products. Product-led discovery, improved documentation and, critically, native cloud delivery are all becoming table stakes for more and more deals.

The way data providers deliver data is changing

Modern delivery programs vary based on the size of the organization and the use cases it supports. Traditionally, data has been delivered either through an owned application (e.g. a terminal) or through bulk file transfer via a legacy protocol like SFTP. In recent years, providers have expanded their options — offering APIs for transactional use cases and pre-built integrations with third-party workflow or business intelligence apps.

Much of the innovation in delivery over the past two decades centered on delivering data through owned or integrated applications. Data companies built successful SaaS businesses and extended their reach through pre-built integrations. Each of these investments succeeded in part because it removed work the data consumer would otherwise have to do to interact with the data. But one form of delivery saw little innovation: "headless delivery," in which analytical data is delivered outside a software experience as tables or files.

Integration is now a strategic barrier to adoption

Headless delivery — often referred to as bulk delivery or data feeds — typically supports analytical use cases where the data is one ingredient in a more complex model. The majority of these deliveries still occur via traditional protocols such as SFTP, placing a heavy burden on the consumer to extract and load the data into their own system and transform it into an analytics-ready state.

That has changed dramatically over the past five years. As data teams moved to the cloud, a new generation of platforms and tools enabled unprecedented scale in analysis, and teams centralized their data into these platforms to run analyses across domains by blending data together into common models. Here's the rub for data providers: traditional headless delivery places a substantial "integration burden" on the buyer — and that burden is now a strategic barrier. Data teams operate at a scale where each additional step creates real complexity, making the extract-and-load process a pain they're no longer willing to accept. Consequently, buyers are demanding that data be delivered directly into these new cloud platforms.

Cloud sharing dramatically reduces the integration burden

Cloud sharing — in which data is made available to a customer natively in their cloud storage or data warehouse — has emerged as an essential part of the delivery toolkit. It's a relatively new phenomenon for two reasons. First, adoption of cloud storage and platforms now represents a majority of buyers. Second, in the past five years each platform has released breakthrough data sharing protocols that allow frictionless sharing of data between accounts. With these protocols, teams can consume a new dataset in their preferred environment without ever extracting it from another server or loading it into their system.

Many platforms have also invested in data marketplaces that let users find, ingest and in some cases buy data products within the platform. Today, marketplaces are less a place where most providers transact and more an increasingly important lead-generation channel for net-new buyers — the ability to explore a dataset before committing to a full trial is a powerful new part of the buyer journey.

Benefits of cloud sharing for data consumers

  • Reduce total cost of ownership by cutting the ELT burden required to consume external datasets and eliminating the need to manage data updates.
  • Improve data quality by eliminating errors that occur in the extract-and-load process.
  • Accelerate time-to-value by making data acquisition as simple as sharing a public identifier.

Why modernizing delivery is critical to data-as-a-service

The ELT burden isn't just a problem for buyers — it's an anchor for data businesses too. Every hour a customer spends loading a provider's data into their system or standardizing file types is an hour not spent finding value in that data. For many providers, cloud delivery is a first step toward a broader strategic vision to scale their operations. Alongside efforts to normalize schemas and improve documentation, native cloud sharing can dramatically accelerate trial and deal cycles that have traditionally limited the velocity of the data business.

"Many of our customers spend too much time on non-value-added tasks to use our data," said Toby Dayton, CEO of LinkUp. "Native cloud sharing allows us to meet our customers where they work — and in doing so, dramatically improve the value our products generate for our customers. Native delivery has allowed us to cut our average trial period from about three weeks to under five days."

Benefits of cloud sharing for data providers

  • Accelerate sales cycles by eliminating the need for customers to extract and load your dataset into their systems, dramatically shortening the average trial and sales cycle.
  • Reach new customers — cloud sharing isn't just a requirement for reaching modern data teams, it's a way to generate new leads by partnering with the destination platforms.
  • Reduce churn. Cloud sharing is increasingly a decision factor, particularly for competitive datasets. Offering it helps deliver value to important accounts and protect against churn.

Principles of a modern delivery program

The goal of a modern delivery program is to meet customers where they work and, in doing so, reduce the time to value for a company's products. Modern programs must meet five key requirements to serve the modern buyer and scale the business:

  • Native — data must be delivered natively into the consumer's cloud data platform; "extract" and "load" should no longer be the customer's responsibility.
  • Agnostic — data must be deliverable to any cloud, platform or destination, since consumers are spread across the major cloud storage providers and data platforms.
  • Secure — the system must be secure and built for non-trusted parties.
  • Scalable — adding customers and clouds should be fast, cheap and predictable.
  • Observable — information about all deliveries should be aggregated and standardized in one place.

5 steps to modernizing your delivery program

Most data providers started exploring cloud sharing only in the past few years, and many are still early in the journey. Here are the phases — and the challenges most companies face — as they move from exploring to piloting to scaling cloud sharing.

1. Get a baseline and set business objectives

Cloud sharing is a strategic capability because it serves a higher goal: building a more scalable data business. Set achievable, realistic goals that tie into your overall strategy. Identify existing investments in cloud sharing, marketplaces or data-as-a-service; evaluate baseline metrics like average sales-cycle length (and trial length specifically), close rate and churn; and map the investment to a leadership priority — whether that's growth or cost reduction.

2. Identify demand for cloud delivery

The most important first step is to identify demand among your existing and future customers. Look at both explicit demand (customers who actively request it) and latent demand (customers who would want it if they knew it was available). Map accounts so your teams know where customers and prospects work and whether they've asked for cloud delivery, and do competitive research — if your competitors already appear in the Amazon, Snowflake and Databricks marketplaces, your customers will likely want those destinations soon.

3. Modernize your delivery stack (build vs. buy)

Delivery teams have traditionally managed bulk deliveries through SFTP, built on decades-old technologies. With cloud sharing, providers need to modernize the stack — or risk burying delivery and engineering teams in custom pipelines that quickly become technical debt. Align delivery priorities against the existing roadmap, secure technical buy-in from product and engineering, and make the build-vs-buy decision. Key challenges to weigh: there's no single "right" destination (buyers span three major clouds and five dominant data platforms); integration is a moving target, as each platform changes its own documentation and protocols; deep domain expertise and commercial relationships are required for each destination; and any new investment must not cost more to stand up and maintain than the value it creates.

4. Go to market with cloud sharing

Sharing data in the cloud isn't just about meeting a requirement — it's about investing in a capability that makes data-as-a-service real. Migrate customers and prospects who are requesting cloud sharing onto the new infrastructure, announce the new capability, build listings on the marketplaces for early discovery and long-tail deals, explore co-selling and co-marketing with the platforms, and experiment with product-led experiences that let prospects interact with your products earlier in the sales cycle.

5. Create data products built for modern delivery

Just as software-as-a-service triggered a revolution in software product design, modern sharing is pushing providers to design data products for the modern buyer — making user experience and customer success more important than ever. Standardize schemas (with data products, the schema is the user interface, so it must be consistent and intuitive), improve documentation so customers can build on your products without constant support, and build in quality so teams can trust the data.

Modernize your delivery stack with Bobsled

Data sharing platforms provide the infrastructure on which modern delivery programs run. A data sharing platform is not another way to share data — it's a single place to provision, orchestrate and manage deliveries across any destination, whether a data warehouse like Snowflake or Databricks, a cloud storage environment like Azure or AWS, or a legacy destination like an SFTP server. A platform like Bobsled should offer three core capabilities for providers:

  • Deliver to any data lake or warehouse. Support delivery to cloud and legacy destinations without leaving the platform where your source data lives — across cloud storage (AWS, Azure, GCP), cloud data warehouses (Snowflake, Databricks, Redshift, BigQuery) and even legacy protocols (SFTP).
  • Build on native sharing capability. Build on the native sharing protocols in each destination to deliver a native experience for consumers — ready-to-query data appears in their environment with no additional integration work required.
  • Expand your cloud delivery program. Bring your entire delivery operation into a single control plane.

With a data sharing platform, you can stand up a new destination in minutes: point Bobsled to where your data lives, pick the bucket or warehouse where you want it to go, and Bobsled handles the rest. There's no need to be an expert in the destination — all you need is your customer's public identifier, with no account to open, bill to manage, or pipeline experts to hire. Supporting delivery into each cloud destination can otherwise take up to six months and substantial investment; with Bobsled, providers deliver in minutes and only pay when they actively share data with customers.

The opportunity ahead

There's never been a better time to be in the data business. Companies now view data as a strategic asset, and investment in external data will only grow as data becomes a key part of every line of business. To meet this opportunity, providers need to think carefully about how they build, share and support data products that ask less of their customers. Cloud sharing and the modernization of delivery are one step toward the bigger vision of data-as-a-service — and if the industry responds, the data business has the opportunity to see extraordinary growth that could parallel the rise of SaaS a decade ago.

Keep reading

More resources

GuidesData Providers

Are Data Marketplaces Worth It?

A few years ago, it seemed like every platform was launching a data marketplace. But today, the enthusiasm has cooled. More and more data and analytics leaders are expressing skepticism. The ones who did succeed approached marketplaces differently.

Read more
SaaSGuides

Playbook: Building a Data Sharing Offering in B2B SaaS

B2B software companies like Stripe, Heap and Braze have turned data sharing into an important part of both their revenue and retention stories. This report will provide everything you need to scope and launch a successful data sharing program.

Read more
Guides

Guide: Sharing and Marketing Data Products on Databricks

This guide will provide everything you need to start sharing data on Databricks whether you sell data as a product or share data with customers as a value-added service. We’ll walk through the basics of the platform and common use cases and then dive into the things you need to know to get started sharing on the Marketplace and beyond.

Read more