Thought Leadership

A Practical Guide to Synthetic Data in Market Research: Part 2

July 15, 2026
Editor’s Note

Hybrid research studies can help extend the value of existing research by combining real human data with synthetic methods. In Part 2 of A Practical Guide to Synthetic Data in Market Research, we explore how Human-Guided Synthetic Research works in practice and why the strongest studies use synthetic approaches to build on human insight—not replace it.

As synthetic data becomes more embedded in real-world research, the challenge is not whether to use it but rather how to apply it to deliver practical value. In our experience, the smartest place to start is a hybrid approach, one that combines traditional research methodologies, real human data and synthetic methods into a single research project. At its core, it starts with real study data, grounded in observed human behavior, which acts as the foundation for everything that follows. Synthetic approaches are then layered in deliberately, not to replace that foundation, but to build on it.

Welcome to Part 2 of A Practical Guide to Synthetic Data in Market Research, where we’ll explore how to design hybrid research studies using synthetic data. This guide builds on the fundamentals covered in Part 1, where we explored the importance of strong data, thoughtful research design and stakeholder alignment. Our goal is to cover how you should think through and design a hybrid study, what you should plan for and expect along the way, and what key factors are essential toward ensuring success.

Imagining our hybrid research study

Synthetic data can quickly become abstract, so this guide uses a simple use case to keep the approach practical. It may not match your exact process, but it should feel familiar enough to apply across a range of research challenges.

First, we’ll define the study:

What’s our real human data? What are we mixing with a synthetic approach to make it a hybrid study? In a perfect world, you might design a hybrid study from the ground up. In reality, most teams aren’t starting there. So, in this example, we’re starting with something much more typical—an existing survey you want to build on.

What’s the research objective? We’ll focus on a research classic: concept testing. In this example, the starting point is a survey you’ve already run to understand current opinions on a product line. That dataset is being used to inform the development of new concepts, which you’ll eventually need to test within the same markets.

Why are you using synthetic data? A clear rationale for using synthetic data is essential, so for this example we’re exploring the core use case mentioned in our first blog: targeting a hard-to-reach audience. In this scenario, some target markets were hard to reach and ended up unevenly represented in your existing data. It was a big enough problem that you never managed to reach the full sample for which you were aiming.

Now you have a familiar problem: limited confidence in some of the audiences that matter the most. And because this work may require multiple rounds of testing, repeatedly returning to the same hard-to-reach audiences will increase costs, extend timelines and risk respondent fatigue.

How does a hybrid study help? Synthetic data’s core strength is allowing you to do more with the data you already have. Some parts of the audience are underrepresented, and at the same time, you need a way to explore and refine concepts without repeatedly going back to the same respondents.

This is where a hybrid approach makes the most sense. You can leverage the data in two key ways:

  • As an anchor for sample augmentation to strengthen coverage within your existing dataset
  • For training conversational persona bots to explore and refine concepts without repeated fieldwork

That decided, you first need to look at your existing data.

"The most robust hybrid studies don't start from scratch. They start with what researchers already know—real human data—and extend its value by using synthetic approaches."

SVP, Head of AI & Secondary Productization

Strengthening the human dataset with sample augmentation

Start with the existing dataset. In this example, the survey has already been completed alongside interviews and other qualitative work, with the goal of understanding how people feel about your current product line. Overall, it’s been well received. You’ve asked the questions you wanted, and the data feels rich, giving you a keen sense of direction.

But there’s a lack of confidence in the parts that matter. Your team has reviewed and flagged that some markets never filled quota, while other groups you realized too late were important went underrepresented. The result is a useful dataset, albeit with areas where you don’t feel able to cut and analyze as robustly as you’d like.

This is exactly where sample augmentation becomes valuable.

It’s the most natural next step here because you’re not trying to create something new—you’re extending what already exists. Augmentation works best when it strengthens groups already present within the dataset, using how those groups relate to the rest of the dataset to build out statistically likely responses.

PRACTICAL EXAMPLE

Imagine your target is a special group of highly engaged brand advocates, those who go beyond typical loyalty and show a deeper connection. This can be a strong candidate for augmentation assuming you have others who have that connection. Their behavior is still rooted in patterns you see across the wider dataset, just at the extreme end. However, if your sample only includes people with some level of brand connection and you try to expand a small group of people with no connection at all, you’re likely to run into issues. Synthetic models learn from relationships already present in the data. Without a meaningful baseline, the model has very little to build from.

To get this synthetic data, you’ll work with experts to run the modeling, but there are a few things you need confidence in before that starts. You’ll need to have reviewed the dataset carefully, identified the groups you want to strengthen and checked that those groups are properly represented in the survey, meaning they’ve engaged with the parts that matter. If their path has lots of skip logic, is too reliant on open ends, or they’ve just not answered reliably, you could face problems with the modeling.

Even if you are working with a supportive vendor, you’ll need to review variables, labels and structure so the data is ready to be used. The model will then learn the relationships within your dataset and generate additional sample, which will need to be reviewed and validated to ensure it reflects real patterns.

Once complete, the model will generate additional rows based on those learned relationships. This gives you a more reliable dataset—one with stronger coverage across the groups that matter, and one that gives you a more reliable base to build from as you move into developing and testing your concepts.

Extending the research with conversational persona bots

With a stronger dataset in place, the next step of our hybrid study can begin. The team goes away to start designing some new concepts based on the feedback. As this is happening, you’re working on the second challenge: how you’re going to explore new concepts with the difficult-to-reach groups.

Recruiting these respondents would be slow and expensive, and showing too many concepts will lead to fatigue. This is where persona bots become valuable.

Persona bots allow you to use the data collected initially in a different way. Instead of treating fieldwork as the finish line, you can use persona bots to interrogate the data, explore needs and sense-check early ideas without going back to the same audience each time.

At this stage, the models are developed with specialist support, who’ll create the personas, but there are a few things you can do to help. Personas can use more than quant survey data. They can draw on richer inputs including how people describe things, the reasoning behind their decisions, and the context behind their behavior. For example, this hypothetical survey is supported by interviews which can be shared with the team creating the personas to allow the models to better reflect people’s language, motivations and lived context.

KEY TAKEAWAY

Adding richness to the personas bots can come in many forms, as long as it’s a relevant trustworthy source. For example, if you wanted more depth and had a tracker of consumer behavior, or even quality syndicated data, you can use that as long as it matches your target groups.

Beyond providing good data, the other task is to clearly set the scope. Personas are most effective when exploring ideas that sit close to the data they are built from. You can test variations and refine concepts, but if you move too far into new territory, outputs will become less reliable.

PRACTICAL EXAMPLE

If your dataset captures reactions to an existing product line, asking about a new product within that same range is a natural extension, and the persona is building on familiar ground. However, if you shift to asking how that product should be sold in-store, you’re moving beyond what the data supports. At that point, the persona stops reflecting observed behavior and starts improvising.

From there, the process moves into building and testing the personas. Inputs are brought together, personas are trained, and outputs are reviewed to ensure they align with the patterns in your data. A human-in-the-loop approach is key here, checking that what you’re seeing is both credible and useful.

At the end of this, you won’t have more data; instead, you’ll have a way to explore and refine concepts before committing to further research and recruitment. The persona bots act as a singular voice representing their segment that you can interact with at your leisure. You can talk to them to help gather insights, and importantly, in our example, they can be used to simulate the reactions of your target groups to new concepts.

This pre-testing and review mean you can whittle down and find the most suitable concepts before you go into any rounds of human testing with these difficult to recruit markets. Even better, the personas will still be useful if none of the concepts resonate and you want to do another round of testing.

"Persona bots are most effective when they help researchers learn and iterate faster, not when they're expected to replace the people they're built to represent."

Senior Specialist, AI & Innovation

The outcome: What a successful hybrid research study looks like

In this hybrid study example: We started with a core survey, strengthened it through augmentation, and extended it through personas.

We have better confidence in the audience we care about. And the data, instead of stopping once fieldwork ends, is now the base on which we can explore ideas, refine concepts and test directions before committing to further research.

This is how hybrid studies become genuinely powerful. There are plenty of other ways to leverage this approach; it not only anchors your work in processes you already understand but also helps ensure you’ve got a human baseline to compare and evaluate the synthetic data against.

This example reflects what we believe is the true value of Human-Guided Synthetic Research. The opportunity is not to replace traditional research methodologies but to extend the value of the research and data you’ve already invested in. When real human data, synthetic methods and expert interpretation work together, organizations can translate insights into action more efficiently while building confidence by remaining grounded in the realities of human behavior.

Key Takeaways: Designing Effective Hybrid Research Studies

  • Hybrid studies generate value by combining traditional methodologies, real human data and synthetic approaches.
  • Human data lays the foundation for the ‘ground truth’ while synthetic approaches help address gaps (strengthening underrepresented audiences via sample augmentation) or extend insights (exploring new concepts and scenarios via persona bots).
  • Human oversight remains essential throughout the modeling, validation and interpretation process.
  • The goal of synthetic data is to extend research value—not replace research itself.

We’re continuing to explore this space every day. If you’re doing the same, it’s worth comparing notes. In the meantime, we invite you to watch “Synthetic Data Without the Hype,” Escalent’s on demand webinar hosted by Chris Barnes and Dyna Boen where they share practical guidance based on what we’ve learned from training our teams and working with F100 clients.


Want to learn more? Let's connect.


Key Questions

1. What is an AI-powered hybrid research study?

An AI-powered hybrid research study combines traditional research methodologies, real human data and AI-driven synthetic approaches within a single research design. Organizations should invest in hybrid studies because they can strengthen underrepresented audiences, support concept development, reduce respondent fatigue and help researchers get more value from existing datasets while maintaining a connection to real human behavior.

2. Why does human judgment still matter in AI-powered hybrid research studies?

AI can help extend, model and explore research data, but human expertise remains critical throughout the process. Researchers determine how synthetic methods should be applied, validate outputs to ensure they remain grounded in real-world behavior, and help take the leap from insightful data to strategic actions. The most effective hybrid studies use AI-powered research techniques to extend the value of human research rather than replace it.

Abhinav Dua
SVP, Head of AI & Secondary Productization

Abhinav is SVP, Head of AI & Secondary Productization. A seasoned strategy consultant and researcher with 16 years of experience, Abhinav has led delivery excellence and consultative insights for Tech-Media-Telecom and Business Process Insights (80+ researchers and consultants) at Escalent. His expertise spans multiple domains (enterprise & consumer tech; media and consumer internet; telecommunications), solutions (unlock growth opportunities; gauge customer pulse; and monetize data/process assets), and methodologies (desk research; social media listening; primary qualitative research; operational analytics; alternative data; and future casting). In his current role, Abhinav is unearthing synergies between human and AI efforts across research and consulting workflows, determining the best use cases and tool bets and ensuring strategic adoption, integration & productization of AI. Simultaneously, Abhinav is productizing existing and conceptualizing new secondary research solutions, while exploring the inevitable cusp of secondary research and AI. Abhinav holds an MBA from the Indian Institute of Management Lucknow, India, and a bachelor’s degree in engineering from NSIT, New Delhi.

James Burchill Headshot
James Burchill
Senior Specialist, AI & Innovation

James Burchill is a senior AI and innovation specialist at C Space, a business unit of Escalent, with ten years of experience in market research. James specializes in evaluating and applying emerging AI tools and methodologies to enhance research outcomes. He holds a Ph.D. in the communication of radical innovations and focuses on translating complex technologies into practical use cases. James often leads training and upskilling initiatives, supporting both internal teams and clients in adopting AI-driven approaches with confidence.