What it actually takes to mine intelligence from your GTM conversations

Want to DIY your conversational intelligence? Here are 8 traps that are the most difficult to build, and where most intelligence tools fall short.

Every revenue team is asking the same question right now:

“We have thousands of call recordings, years of email threads, a CRM full of deal history. Why don't we just point Claude at them and extract the insights ourselves?

We sat down with Octave’s technical team to understand why. What makes Claude, RAG, call recorders like Gong, and DIY solutions get it so wrong?

We plan to share a deep dive, and until then, here’s a preview of 8 parts of Octave that are the most difficult to build.

These contribute to the bigger engine that generates intelligence and agent context for more precise and accurate insights and content.

Before we dive into them, we need to acknowledge the most difficult ingredient to DIY.

Time.

Our team has been — and we say this with affection — unreasonably tenacious at whack-a-mole’ing edge cases over the last 4 years. 

Octave specializes in GTM for B2B companies, so we’ve gone deep into making our findings correct. To get reliable answers, the team had to build something, then wait for a deal to come along that would break it with a detail they’ve never seen before. Then, they could design new logic to accommodate it. After that, they had to wait for more deals to confirm those design decisions. And repeat. 

This is why building truly precise and accurate GTM intelligence tools cannot be done overnight. You can’t compress four years of iteration cycles. 

Learning cycles are tethered to the timing of your deals and market, not to how fast you can build. Building is the fun, fast, and easy part — but we talk to so many teams who have gotten pretty far with DIY, only to realize it doesn’t hold up over time.

The cost of getting it wrong

What’s dangerous about most intelligence tools, whether DIY or off the shelf, is that they give incorrect answers that look right on the surface.

For example, a Gong sales call gets an “A” grade because the customer’s sentiment is positive, and the talk time and question counts look good.

But when you really look under the hood, it turns out that the rep failed to bring up the standard kill-shot for a common objection. Or they showed a financial services case study to a healthcare buyer. This call is actually a miss. The tool didn’t know what “good” really looks like. 

(Octave tackles this by comparing conversation findings to your Octave library.)

Another issue is that most tools tend to look for what they know.

Imagine you ask Claude to find objections in your last 20 call transcripts. 

It may skew toward finding pricing objections, either because that’s an obvious thing to look for, or because you’ve programmed it to be especially sensitive to that. But you might not realize that customers are describing security concerns with phrases you don’t realize you should scan for. 

If security objections are what’s actually killing your deals, you may never know.

(Octave’s approach is to offer a wide range of top-down annotations, as a catch-all extractor that finds the unknown unknowns.)

In our next technical deep dive, we’ll share specifics about how Octave solves these problems and more. For now, here are a few of the things that are hardest to DIY when you build an intelligence tool, ranging from “annoyance” to “technical hornet’s nest.”

Ingesting source material: Medium difficulty

Octave extracts insights across many sources, including meetings, emails (sent/replies), CRM changes (opportunities created, deals won, deals lost), LinkedIn messages, ads, and product telemetry from your data warehouse.

It also indexes (and refreshes) resources like your website, Notion, Google Drive, Linear, GitHub. 

This means we’re integrating with Fathom, Gong, Salesforce, HubSpot, Notion, Google, Linear, Snowflake, Databricks, GitHub, and so many more (via API, webhook, MCP, etc.). 

It’s not technically difficult. But do you really want to maintain those integrations?

All we’re really doing is staying on top of the format of data coming in: we standardize the messy upstream inputs into something that our downstream engines can process consistently.

But they break constantly, because vendors change their data schemas and API’s deprecate. It takes a lot of opinions, testing, and seeing it break a lot.

Final score: Pain in the butt, but doable.

Scenario understanding: Medium-to-hard difficulty

A transcript arrives. Is this a sales call or a support escalation? A prospect or an existing customer? Net new or upsell? Which product?

Getting the scenario wrong muddies everything downstream. 

People commonly assume that a call just attaches to an opportunity in your CRM, and that’s enough context to discern the scenario. But in practice, CRMs are rarely clean enough to support this.

Octave has built the logic to detect what’s what. For example, we filter out internal calls, like a bouncer that turns away meetings that aren’t eligible for analysis. Or if someone on a call mentions a certain phrase, we drop it.

We created this logic by hand-building it over time. Half of that comes from the narrow domain specialty of B2B GTM, and half comes from referencing your company’s Octave library to increase the bouncer’s comprehension. Your GTM ontology is like a lens with the right prescription to see better.

Speaker attribution: Medium difficulty

Knowing who’s speaking is very important, not just because you want to understand a specific deal, but because you want to see patterns in your GTM.

If you know their title is CMO, then you can start detecting patterns across all CMO’s. If you know their company too, then you can start saying, “executives at companies with over $100M in funding tend to worry about X.”

So if your call transcript only gives you the speaker’s name or email, how do you know which Mary at Microsoft is mary@microsoft.com? Octave maintains a database of people and companies to quickly resolve identities. This includes firmographic and funding data, which helps match companies to segments.

This isn’t super challenging to build, but the data expense can add up, and making mistakes has a big impact.

Annotating findings: Extremely difficult

This is where it gets fun. We tag every call, quote, conversation with extracted findings and the entities in your library they match to.

Octave casts a very wide net to annotate snippets of call transcripts (and other conversations). We maintain a quickly growing set of over 150 extraction types that will find and label meaningful phrases. 

Octave’s extractors are opinionated about what we’re looking for, which helps with correct answers. Again, this is informed by a combination of (1) our years of building B2B GTM-specific extractors and (2) your specific company’s GTM ontology in your Octave library.

This means that your specific personas, use cases, and product information inform how the extractors behave to produce correct results.

We tag findings with information like the speaker, their email, a persona ID, a company domain, the segment IDs, deal outcome, opportunity ID (like HubSpot ID), an event ID (for example, a specific call), sentiment, window of the call, entity IDs (if the snippet is about a competitor, use case, objection), the literal quote snippet, and more.

This opens up a whole universe of questions to ask with reliable answers:

  • Find all the hesitations on our deal with Acme Co.
  • Find everything that CTOs and CISOs say about security reviews
  • Find anything related to switching costs and rip-and-replace, organized by size of company talking about it
  • Figure out why customers switched away from Competitor A to us, divided by persona

Maintaining these takes work, and that’s just for the first-time labeling. What happens when you change something about your strategy? Don’t you want to go back and re-label past calls to understand them in a new light? This brings us to our next point…

Updating call transcript annotations: Difficult

If you add a new competitor to your strategy, or launch a new product use case, you now need to go back and update the transcript annotations that you’ve processed in the past. 

Otherwise, you can’t ask, “look for companies that would benefit from this new use case” and see historical answers.

Octave updates transcript annotations automatically when you make a change to your Octave library, retroactively extracting new intelligence based on information you didn’t know about at the time of originally processing the transcript.

GTM ontology (Octave Library): Medium difficulty

Creating a library of your GTM strategy (or type 2 context) is not a technically difficult task. However, it’s a bundle of opinions that have an effect on returning precise and accurate answers. For example: 

  • What counts as a persona, and what doesn’t?
  • What are its edges?
  • Why is this persona subordinate to that segment, but not another?
  • How do you know when your product feature “solves” a common problem for that persona — versus just “relates” to it?

We’ve baked years of our accumulated judgment into these opinions. Building the container is the easier part, filling it is the harder work.

Reconciling call insights with Library: Difficult

Octave’s secret sauce is that it compares your library with the conversations that are happening. Your library is like your hypothesis of what your GTM strategy should be, while the conversations are reality. Octave is always checking one against the other. 

So every time Octave finds meaning in a quote from a customer, it checks it against what you already believe. If it finds something there — oh, this CISO is talking about a pain that appears on the CISO persona entity — then it links that pain with the library entities for that pain, that persona, the segment the CISO is in, etc. 

And the matches are not always clean and obvious. Octave not only tags a finding with a matched entity, but also with the closest match entity. 

It continues to listen and, maybe, notice that CISOs keep speaking about a pain that’s not in your library yet — then create a suggestion. We’ll share more about suggestions and learnings in a deeper dive post.

Connective tissue: Extremely difficult

The hardest part of all this is to keep all these machines running together. 

Hundreds or thousands of findings stream in from the 10 calls your sales reps had in the last 15 minutes alone. Octave is annotating them, attributing them to speakers, and then checking them against the library. 

If something exists, it’s linking the entity, and if it doesn’t, it’s putting the finding into a holding pen. Batch jobs are deciding, right now, whether the accumulating evidence would warrant a suggested update to the library to fill in that gap. While the decision’s being made, five more calls are landing — their findings are strengthening the case, or not. 

How do you manage this? Do you have look-back windows to help with that decision making?

And then imagine that you update an entity in your library: your competitor announces a new feature, so you add a line to your Library entity for that competitor. This change cascades to all the other entities it touches — personas, playbooks, use cases, pains, and more.

There’s a Cambrian explosion of relationships to maintain. One deal in your CRM is attached to the findings database through every annotated snippet from calls and emails with that company. It’s also attached to library entities (segment, personas, use cases, objections they raised, etc). It’s attached to the resolution layer — the real-world identity data that made the persona-and-segment matches possible. 

A deal is also attached to the outside world: the market that company plays in. So Octave is built on two context streams: 

  • Internal stream: everything your business generates, like calls, emails, deals, documents, etc.
  • External stream: news from the market, competitors, customers, their markets, industry shifts, micro and macro trends, etc.

The learning loops for each of these are constantly checking against your library: does your hypothesis hold up to reality? How might it need to change? How is it spot on? 

And every accepted suggestion to update the library updates what the extractors look for in your call intelligence — the next calls are read more sharply, and all the calls that came before it.

Nothing in the graph is static, and that’s the machine you’d be signing up to build.

This is just a peek of what’s under the hood of Octave. If you have questions for our technical team, please get in touch — we love hearing how people have been figuring out DIY!

The foundation for agentic GTM

Placeholder Image