TL;DR: Sales forecasting is a data engineering problem first. A model can only reason about the deals it can see, tied to the right account, the right people, and the right opportunity. At GTM Engine, the forecast sits on five layers: a CRM sync you can audit, identity resolution, an activity graph built from email, calendar, and call data, typed AI fields with history, and alerts that fire when risk changes. Every forecast number should trace back to the evidence that moved it.
Ask an AI model to forecast your quarter from CRM fields and you get a confident summary of stale data. AI sales forecasting is only as accurate as the records underneath it, and CRM records are often a lagging, subjective snapshot of whatever reps remembered to update.
When we started building forecasting at GTM Engine, we did not start with a dashboard. We started with a question: what would the data infrastructure need to look like if every customer interaction could update the forecast automatically?
Our answer is a revenue evidence graph. A revenue evidence graph is the connected set of CRM records, emails, meetings, calls, transcripts, participants, and associations that describe what is happening in a deal, plus rules for who owns each field and a history of every change. The forecast is the last thing built on top of it.
Forecasts fail because the CRM is a lagging, subjective snapshot
Your forecast call reviews what reps typed, and reps type late.
A CRM forecast reflects the last time someone touched a field. Close dates move after the deal has already slipped. A new stakeholder joins the email thread, but nobody adds them to the opportunity. A meeting happens and lands on the account instead of the deal. A buyer mentions on a recorded call that procurement is backed up until next quarter, and that sentence stays trapped in a transcript.
By the time the forecast call starts, categories and health scores are opinions. Adding AI on top of those opinions produces faster opinions.
The fix is upstream. Before a model can tell you whether a deal is real, the system has to know what happened, who was involved, and which deal it belongs to.
AI sales forecasting starts with a CRM sync you can audit
If you can't say where a field value came from, you can't trust the forecast built on it.
AI sales forecasting starts with treating the CRM as an event source whose changes are captured, ordered, normalized, and applied safely. We do not treat Salesforce or HubSpot as an API we call occasionally.
Inbound CRM changes land in a sync ledger before they touch core records. Each ledger entry keeps the CRM ID, the target table, the raw record payload, a delete flag, the sync mode, timestamps for when the record was updated and when it was synced, a retry count, and the source. A drain process works through pending entries in a deterministic order, batches account and contact updates, handles associations.
The second half of trust is field ownership. Each mapped field stores its CRM property ID, data type, enum values, record type, whether the CRM allows edits, and a sync direction: from the CRM, to the CRM, or bidirectional. Inbound transforms only accept from-CRM and bidirectional fields. Outbound writes only send to-CRM and bidirectional fields, through a single provider-agnostic router that detects Salesforce or HubSpot and checks the organization's sync setting (up, down, both, or off) before writing. When an admin maps a field the CRM marks read-only, we downgrade it at save time, so GTM Engine never pretends it can write a value the CRM will reject.
Forecast accuracy starts with knowing who owns every field.
Identity resolution comes before prediction
Duplicate accounts and contacts quietly split or inflate your pipeline.
Identity resolution is the process of deciding which account, which people, and which opportunity a piece of evidence belongs to. It has to happen before any model sees the data.
GTM Engine matches accounts by CRM ID, then internal ID, then canonical domain. Contacts match by CRM ID or canonical email. Opportunities match by CRM ID and internal ID, with conservative fallback logic.
The goal is unglamorous: one Acme account instead of three, so every rollup counts each deal once.
AI cannot forecast on top of duplicate identity.
Every buyer interaction attaches to a specific deal
A meeting that lands on the wrong record is a signal your forecast never sees.
Every email, meeting, and call becomes an activity in one normalized structure. An activity stores its source and source ID, a unique ID, the raw data, the owner, direction, timestamp, status, interaction type, and a flag for whether external participants were involved. Join tables link each activity to contacts, opportunities, and users. Transcripts are first-class records with normalized text, the raw transcript JSON, participants, call metadata, recording links, and a link back to the activity.
Participants come from wherever the conversation happened: from, to, and cc on email; attendees on Google and Microsoft calendar events; and participant lists from call recorders including Gong, Chorus, Fireflies, Fathom,
Read.ai, Circleback, Sybill, and Recall. A shared matching path separates internal users from external buyers.
When a new external email address shows up on a relevant thread or meeting, GTM Engine can create the missing account first, then the contact, then associate both to the activity and the opportunity. It only creates records when there is enough selling context, which keeps noise out of the CRM.
Association writes are idempotent. They use primary keys and conflict-safe inserts, so processing the same meeting twice produces each relationship once. A single pass for calendar and call activity can link the activity to users, contacts, and opportunities, and link contacts to accounts and opportunities. Contact-to-account and contact-to-opportunity associations can sync back to the CRM, so the new stakeholder shows up there too.
A new buyer joining a deal thread is a forecast event.
AI output lands in typed fields with history
A paragraph in a chat window can't be filtered, trended, or alerted on.
Typed persistence is the line between an AI demo and an AI system. When a model produces a forecast date, a health score, or a promoter score, that output has to land in a structured field with history and downstream behavior.
GTM Engine's AI prompt tasks return structured objects. Each output is validated against a schema, retried if validation fails, and mapped into typed database columns or JSON fields. The standard set covers all three core records:
Account — AI-generated fields: Propensity signals, overall Propensity To Buy score, technographic and firmographic data, product offerings, competitors, go-to-market strategy, account Skill file
Opportunity — AI-generated fields: Health score and reasoning, AI forecast close date and reasoning, deal gaps, suggested strategies, overall interest score, forecast category, path to close, predicted stage, deal participants, methodology progress
Contact — AI-generated fields: Interest level, contact role, promoter score, ICP persona, reasoning fields
The standard set is a starting point. Teams can define their own AI fields in any supported data type: text, number, date, picklist, dynamic picklist, boolean, JSON, user, and others. A custom field gets the same treatment as a built-in one: a schema the output has to pass, a typed column, and history. If your methodology needs a "procurement path confirmed" checkbox or a "competitor mentioned" picklist, the model fills it from deal evidence, and you can filter, trend, and report on it like any other field. Custom AI fields sync to Salesforce and HubSpot through the same field mapping and sync direction rules as every other field, so the value your reps see in the CRM matches the one the forecast uses.
Accounts get two layers of context. The first is research and enrichment built from first-party and third-party data: propensity signals that roll up into an overall
Propensity To Buy score, technographic and firmographic data, the company's product offerings, its competitors, and its go-to-market strategy. Anything else you want to track at the account level becomes a custom field.
The second layer is the account Skill file. A Skill file is the durable knowledge your team gathers while selling to a company: its priorities, procurement process, budget cycles, and how decisions get made. Teams configure what it tracks.
Genie, GTM Engine's AI agent, reads the Skill file when it builds meeting prep, helps a rep strategize through a deal, and works on upsells and renewals. The next deal at that account starts with what your team already learned on the last one.
Because these are fields, the same outputs power forecast tables, filters, history comparisons, dashboards, notifications, and answers from Genie.
Deal health combines engagement quality with timing
A green deal with a forecast date outside the quarter is still a risk to this quarter.
Deal health in GTM Engine has a strict definition. Health scores default to a 1 to 10 rubric. A deal counts as healthy only when it has health data, scores 7 or higher, and its AI forecast close date stays inside the expected quarter.
Deals without health data are marked unknown rather than at risk, because missing evidence and bad evidence call for different actions. At-risk logic combines the health score, the rep's close date, the AI forecast close date, current quarter boundaries, and closed status. A deal the buyer loves that won't sign until next quarter shows up as a timing risk in this quarter's number.
A forecast date is only as useful as the reasoning behind it
Your CRO will ask why the date moved. The system should already know.
Opportunity reevaluation runs as an org-level workflow. It weighs the opportunity context, recent activity, health reasoning, deal gaps, the prior forecast, the AE's close date, your stage definitions, and methodology progress. It stores both the AI forecast close date and the reasoning that produced it. When the date changes, that update can trigger notifications and write to history.
The date is the conclusion. The reasoning is what a manager can actually argue with in a pipeline review.
Deals deteriorate in the people graph first
The amount field is the last place a dying deal shows up.
Contact-level signals are leading indicators. A champion goes quiet, an economic buyer stops attending meetings, an evaluator stalls, and the amount and stage fields stay exactly where they were.
GTM Engine tracks four AI signals per contact. Interest level captures engagement strength. Promoter score separates advocates from passive participants. Contact role classifies each person as economic buyer, decision maker, evaluator, technical buyer, executive sponsor, influencer, or another role. ICP persona classifies the person by function, seniority, engagement pattern, and profile. Contact AI history records changes to these fields over time, so you can see when a champion started cooling off.
Forecast history and risk alerts close the loop
A rollup gives you the number. History tells you why it moved.
Forecast history reconstructs past opportunity snapshots and compares them to the current state across amount, close date, stage, forecast category, health score, and AI forecast close date. Changes fall into movement buckets: entered forecast, exited forecast, became healthy, became at risk, and amount changed. The history tab on the Forecast page answers the question every forecast call starts with: what changed, and why? Forecast Map shows the same movement as category flow and risk flow.
Risk should not wait for the next forecast call either. Notification events fire when a deal moves out of the quarter, when the AI forecast date moves out of the quarter, when a deal becomes unhealthy, and when a close date comes due. Each notification carries the owner, account, opportunity, previous and current dates, previous and current health score, the Slack channel, and a deep link to the record. Workflow-driven opportunity updates and record-updated triggers can both publish these events.
What this architecture does not do
Trust also depends on being clear about the limits.
Sync is near-real-time. A change in Salesforce or HubSpot arrives through sync jobs and webhooks, so there is a short lag before it reaches the forecast.
Health is unknown until there is evidence. A brand-new opportunity with no activity will not get a health score by default, and we treat that as a gap to fill instead of guessing.
Record creation is deliberately conservative. Because GTM Engine waits for selling context before creating accounts and contacts, a legitimate new stakeholder can take a little longer to appear than a raw email import would.
Reasoning follows your definitions. Reevaluation reads your stage definitions and methodology progress, so vague stage definitions produce vague reasoning.
The forecast is the output of a loop
Dashboards are easy. The loop underneath them is the product.
Every CRM and communication event gets ingested, resolved to the right identity, associated to the right contacts, accounts, users, and opportunities, and turned into structured AI fields with reasoning and history. Movement gets detected, and the owner hears about it.
Forecasting becomes trustworthy when every number can be traced back to the evidence that changed it.
Frequently asked questions about AI forecasting data
What is a revenue evidence graph? A revenue evidence graph is the connected set of CRM records, emails, meetings, calls, transcripts, and participants tied to specific accounts, contacts, and opportunities, with field ownership rules and change history. It is the data layer an AI forecast reasons over.
Why can't AI forecast accurately from CRM fields alone? CRM fields reflect what reps updated, not what happened. Without activity data and identity resolution, a model is summarizing stale and sometimes duplicated records.
Does GTM Engine sync with the CRM in real time? Sync is near-real-time. It combines frequent sync jobs, webhooks, a sync ledger, and UI and API refresh, and it prioritizes ordered, complete changes over instant ones.
Which CRMs does GTM Engine support for two-way sync? Salesforce and HubSpot. Outbound writes go through one router that detects the provider and respects each field's sync direction.
How does GTM Engine avoid duplicate accounts and contacts? Accounts match by CRM ID, internal ID, then canonical domain. Contacts match by CRM ID or canonical email. New records are only created when there is enough selling context.