← Home

Deal Appetite Experiment

From a latent sales-ranking conjecture to stateful process prediction.

Companion repository: vbnovikov/deal-appetite-experiment

After looking at CRM data and business workflows across several clients, I started to wonder whether there was some underlying combination of factors that could reliably predict deal closing probability better than a simple lead score.

Things like number of touchpoints, response patterns, friction, time between interactions, stage movement, and other parts of the customer journey all seemed potentially useful when considered together.

That led to the idea that there may be a generalized ranking function over opportunities, with some part of the outcome determined by unobservable variation. If that were true, the useful question would be whether the recoverable part of that ranking could be estimated from observed CRM data, even when the full underlying process could not be observed directly.

Short Version

The original conjecture did not survive in its broadest form.

Synthetic data showed that a recoverable appetite ranking can be learned when the structure is explicitly built into the data-generating process. Real CRM opportunity fields, early interaction aggregates, and surface-level interaction text did not provide strong or stable out-of-sample ranking of terminal outcomes.

The more useful result appeared after changing the target from distant outcome prediction to process-local prediction.

On a real event log, current state alone predicted the next event with about 0.6660.666 accuracy. Adding process history increased accuracy to about 0.7810.781. Adding resource identity increased it to about 0.8380.838. Adding simple graph-derived relational context increased it again to about 0.8460.846.

So the final claim is narrower:

Business state may be better represented for decision-making as a history of connected entities and events than as the current values of those entities alone.

Hypothesis

H1H_1

There exists a recoverable ordering function over observable CRM variables:

Ai=f(xi)+εiA_i = f(x_i) + \varepsilon_i

such that f(xi)f(x_i) preserves useful relative ordering even when εi\varepsilon_i is unobserved.

Here, xix_i means the feature vector for opportunity ii: the list of measurements the model is allowed to see. Plain-language explanations of notation, feature vectors, feature groups, and preprocessing are included in Appendix A1, Appendix A2, Appendix A3, and Appendix A6.

H0H_0

Observable CRM variables do not recover a stable relative ordering of opportunities beyond noise and unobserved variation.

Experiment Setup

I generated 5,000 synthetic opportunities with four observable variables:

  • fit
  • urgency
  • trust
  • friction

Each variable was sampled independently from a standard normal distribution:

fi,ui,ti,ri∼N(0,1)f_i, u_i, t_i, r_i \sim \mathcal{N}(0,1)

The observable component of appetite was defined as:

Si=1.2fi+0.9ui+0.7ti−1.0riS_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i

I then added an unobserved term:

εi∼N(0,1.52)\varepsilon_i \sim \mathcal{N}(0,1.5^2)

giving total latent appetite:

Ai=Si+εiA_i = S_i + \varepsilon_i

AiA_i was converted to a closing probability using the logistic function:

pi=11+e−Aip_i = \frac{1}{1 + e^{-A_i}}

and the observed outcome was sampled as:

Yi∼Bernoulli⁡(pi)Y_i \sim \operatorname{Bernoulli}(p_i)

A logistic regression was then trained using only fif_i, uiu_i, tit_i, and rir_i. The latent appetite, unobserved term, and true closing probability were withheld from the model.

The point of the synthetic setup is to separate what is observable from what is hidden. Appendix A2 explains how observable columns become model inputs, and Appendix B2 and Appendix B3 explain the ranking metrics used to evaluate the model.

Evaluation

The model was evaluated against three different targets.

The metrics used here are defined in plain language in Appendix B2 and Appendix B3.

1. Observed outcome

How well does the model rank deals that actually closed above deals that did not?

This is measured using ROC AUC against YiY_i.

2. Total latent appetite

How closely does the model's ranking agree with the full latent quantity:

Ai=Si+εiA_i = S_i + \varepsilon_i

This includes the unobserved term, so perfect recovery should not be possible.

3. Systematic appetite

How closely does the model recover the ordering generated by the observable component:

Si=1.2fi+0.9ui+0.7ti−1.0riS_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i

This is the part of the synthetic process that should, in principle, be recoverable from the available data.

Results

The model achieved an ROC AUC of approximately:

AUC⁡≈0.811\operatorname{AUC} \approx 0.811

against the observed close/no-close outcome.

That was encouraging, but the more important comparison was against the quantities that are known only because the data is synthetic.

The model had limited agreement with total latent appetite,

Ai=Si+εiA_i = S_i + \varepsilon_i

because εi\varepsilon_i is deliberately hidden:

ρs(p^,A)≈0.791\rho_s(\hat{p}, A) \approx 0.791

Its ranking of the systematic component SiS_i, however, was nearly perfect:

ρs(p^,S)≈0.999\rho_s(\hat{p}, S) \approx 0.999

In the synthetic setting, the recoverable part of the ranking function was therefore recoverable from the observable variables.

The Problem

The synthetic experiment does not establish that a comparable latent variable exists in real CRM data.

The systematic component was defined directly from the same variables given to the model:

Si=1.2fi+0.9ui+0.7ti−1.0riS_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i

A model recovering that ordering therefore shows that the estimation procedure works under the assumptions of the simulation. It does not show that a comparable one-dimensional quantity exists in real CRM data.

The unobserved term creates another problem. In a real business process, εi\varepsilon_i is not a clean random variable whose distribution is known. It can contain missing interactions, individual behavior, timing, relationships, information outside the CRM, and variables that were never recorded.

At that point, the conjecture loses a clear interpretation because the unobserved component can contain almost anything the model does not capture.

Real CRM Data

The next experiment used a real CRM sales dataset with resolved opportunities linked to account, product, and sales-team data.

The model only used information that would have been known while the opportunity was still open. Fields like the final close date or realized deal value were excluded because they would leak the outcome into the model and make the test meaningless.

To avoid leaking future information into the training set, the data was split chronologically:

Train={i:ti<T}\text{Train} = \{i : t_i < T\} Test={i:ti≥T}\text{Test} = \{i : t_i \geq T\}

The model was trained on earlier opportunities and evaluated on later ones.

This is a temporal holdout: the test cases occur later in time than the training cases. Appendix A7 explains why the same boundary also matters for preprocessing.

The baseline result was:

AUC⁡≈0.526\operatorname{AUC} \approx 0.526

which is only slightly above random ranking.

Feature Ablation

I then removed and added groups of variables while keeping the model and temporal split fixed.

The results were:

AUC⁡Core≈0.530\operatorname{AUC}_{\text{Core}} \approx 0.530 AUC⁡Core + Rep≈0.540\operatorname{AUC}_{\text{Core + Rep}} \approx 0.540 AUC⁡Core + Rep + Account≈0.525\operatorname{AUC}_{\text{Core + Rep + Account}} \approx 0.525

Salesperson and team information added a small amount of predictive value. Adding account identity reduced performance on the future holdout set.

The static CRM fields available in this dataset therefore provided very little stable out-of-sample ranking power.

Interaction History

The weak performance of the static CRM model raised a more specific question: does the history of interactions with an opportunity contain useful information that is missing from the opportunity record itself?

The next experiment began with a comparison between static CRM data and static CRM data plus interaction-derived information. In practice, the available interaction records introduced a narrower test: whether coarse first-30-day activity summaries and surface-level interaction text contained stable ranking signal under temporal validation.

Static CRM Model

The first model estimates the probability that opportunity ii will be won using only static CRM features:

p^i(s)=P(Yi=1∣xi(s))\hat{p}_i^{(s)} = P(Y_i = 1 \mid \mathbf{x}_i^{(s)})

where xi(s)\mathbf{x}_i^{(s)} is the feature vector containing the static CRM information supplied to the model.

If this notation is unfamiliar, read xi(s)\mathbf{x}_i^{(s)} as "the static-data row for opportunity ii after it has been converted into numbers." Appendix A1 explains the notation, and Appendix A2 gives concrete examples.

The static feature set includes fields such as:

  • product
  • account
  • sector
  • salesperson
  • regional office
  • company revenue
  • number of employees
  • deal value
  • engagement month

Categorical fields are one-hot encoded and numerical fields are standardized using statistics from the training set. The preprocessing details are included in Appendix A6 and Appendix A7.

Interaction-Derived Features

A natural next model would include information derived from the communication history associated with the opportunity:

p^i(s+b)=P(Yi=1∣xi(s),xi(b))\hat{p}_i^{(s+b)} = P\left( Y_i = 1 \mid \mathbf{x}_i^{(s)}, \mathbf{x}_i^{(b)} \right)

The new term,

xi(b)\mathbf{x}_i^{(b)}

is another feature vector, this time derived from interactions rather than static CRM fields.

The superscripts here label feature groups. They are not exponents. xi(s)\mathbf{x}_i^{(s)} means static features, and xi(b)\mathbf{x}_i^{(b)} means behavioral or interaction-derived features.

Depending on the interaction record, these features can describe things such as:

  • number of interactions
  • time between interactions
  • recency of communication
  • direction of communication
  • who initiated the interaction
  • interaction type
  • properties extracted from communication text

For example, an opportunity might have the following interaction history:

Day 1   Salesperson → Customer
Day 3   Customer → Salesperson
Day 8   Salesperson → Customer
Day 9   Customer → Salesperson

That sequence can be converted into numerical features such as:

interaction_count       = 4
customer_replies        = 2
salesperson_messages    = 2
days_since_last_contact = 1
mean_response_delay     = ...

Those values can then be encoded into xi(b)\mathbf{x}_i^{(b)} and supplied to the model alongside the static feature vector.

The combined expression can therefore be read as:

the probability that opportunity ii is won, given both its static CRM data and its interaction history.

The motivating comparison is whether the combined representation improves on static CRM data alone:

AUC⁡s+b>AUC⁡s\operatorname{AUC}_{s+b} > \operatorname{AUC}_{s}

However, the available interaction records were not clean opportunity-level histories. Many interactions were better interpreted as relationship-level observations between salesperson and customer contact. The implemented baselines therefore tested a narrower question first: whether first-30-day activity summaries or first-30-day lexical content contained stable ranking signal under temporal validation.

Implementation details for preprocessing, modeling, and evaluation are included in the appendices, especially Appendix A6, Appendix A7, and Appendix B1.

Interaction History Results

The implemented interaction-history baselines tested:

AUC⁡activity\operatorname{AUC}_{\text{activity}}

for first-30-day activity summaries, and:

AUC⁡text\operatorname{AUC}_{\text{text}}

for first-30-day TF-IDF text features.

The initial interaction baselines did not find useful out-of-sample ranking signal. The activity-only temporal holdout produced:

AUC⁡activity≈0.473\operatorname{AUC}_{\text{activity}} \approx 0.473

and the text-only baseline produced:

AUC⁡text≈0.495\operatorname{AUC}_{\text{text}} \approx 0.495

Both are at or below the random-ranking benchmark:

AUC⁡=0.5\operatorname{AUC} = 0.5

This does not imply that interaction history is irrelevant. It shows that these initial representations did not recover stable outcome-ranking information in this dataset.

One likely issue is compression. Interaction history contains more detail than the static opportunity record, but much of that detail is lost before it reaches the model.

For example, a sequence such as:

Day 1   Salesperson → Customer
Day 2   Customer → Salesperson
Day 3   Customer → Salesperson
Day 12  Salesperson → Customer

might be reduced to features such as:

interaction_count       = 4
customer_messages       = 2
salesperson_messages    = 2
mean_response_delay     = ...
days_since_last_contact = ...

Those aggregate values preserve some information about the interaction history, but they no longer describe the sequence itself.

Two opportunities can therefore have similar aggregate features while having very different histories.

For example:

Opportunity A

Salesperson → Customer
Customer → Salesperson
Salesperson → Customer
Customer → Salesperson

and:

Opportunity B

Salesperson → Customer
Salesperson → Customer
Customer → Salesperson
Customer → Salesperson

can produce similar counts even though the order of events is different.

The same issue applies to timing. Four interactions spread evenly across a week describe a different process from four interactions followed by a month of silence, even if both opportunities have the same total interaction count.

This led to the next question: Is the useful information contained in the trajectory of an opportunity rather than in aggregate features calculated from that trajectory?

Sequential Customer Journeys

The interaction-history experiment reduced a sequence of events to summary variables such as interaction count, recency, and response timing. That loses information about the order in which those events occurred.

The next experiment tested whether preserving more of the sequence improves prediction.

Instead of treating an opportunity as one static feature vector, each customer journey is represented as an ordered series of events.

This is the sequence version of the same idea. A feature vector stores a fixed list of inputs for one example; a sequence stores an ordered list of events for one example. Appendix A2 gives the feature-vector version first, and Appendix A5 shows how histories are converted into model inputs.

For session ii, let:

Ei=(ei1,ei2,…,eini)E_i = \left( e_{i1}, e_{i2}, \ldots, e_{in_i} \right)

where:

  • EiE_i is the complete observed journey for session ii
  • ei1e_{i1} is the first event
  • ei2e_{i2} is the second event
  • nin_i is the total number of events observed in that session

A simple journey might look like:

home
→ product_page
→ cart
→ checkout
→ confirmation

The order is important because a customer who reaches checkout has followed a different process from someone whose session ends on the product page, even if both generated the same number of events.

Journey Prefixes

The model should only use information that would have been available at the time a prediction was made.

For that reason, the experiment does not give the model the complete journey at once. It constructs progressively longer prefixes of the journey.

The first kk events are represented as:

Ei(k)=(ei1,ei2,…,eik)E_i^{(k)} = \left( e_{i1}, e_{i2}, \ldots, e_{ik} \right)

For example:

k = 1

home
k = 2

home
→ product_page
k = 3

home
→ product_page
→ cart
k = 4

home
→ product_page
→ cart
→ checkout

At each point, the model estimates:

p^i(k)=P(Yi=1∣Ei(k))\hat{p}_i^{(k)} = P \left( Y_i = 1 \mid E_i^{(k)} \right)

Here:

  • Yi=1Y_i = 1 means that session ii eventually results in a purchase
  • Ei(k)E_i^{(k)} is everything observed in that session up to event kk
  • p^i(k)\hat{p}_i^{(k)} is the estimated probability of eventual conversion based only on that partial journey

The expression can be read as:

the estimated probability that session ii eventually converts, given the first kk events observed in its customer journey.

The experiment then compares:

AUC⁡(1),AUC⁡(2),AUC⁡(3),AUC⁡(4)\operatorname{AUC}(1), \operatorname{AUC}(2), \operatorname{AUC}(3), \operatorname{AUC}(4)

If the ordering becomes more accurate as kk increases, additional process history is providing useful information about the eventual outcome.

Avoiding Outcome Leakage

The dataset contains a Purchased field on every row in a session.

That creates an important problem:

If a session eventually results in a purchase, earlier rows in the same session can also contain:

Purchased = 1

even though the purchase had not happened yet.

Using that field as an input would effectively tell the model the answer in advance.

The confirmation event creates a similar issue. Reaching confirmation means the purchase has already occurred, so including it in the predictive history would make the experiment trivial.

For that reason:

  • Purchased is used only as the final target
  • confirmation is treated as a post-outcome event
  • only events that occur before confirmation are used as model inputs

This keeps the question meaningful: given only what had happened so far, could the model rank sessions by their eventual probability of conversion?

This is the same leakage rule used throughout the experiments: do not let information from after the prediction time enter the inputs. Appendix A7 describes the rule in preprocessing terms.

Constructing the Sequences

The event data is first ordered by session and timestamp:

customer_journey["Timestamp"] = pd.to_datetime(
    customer_journey["Timestamp"]
)

customer_journey = (
    customer_journey
    .sort_values(["SessionID", "Timestamp"])
    .reset_index(drop=True)
)

Each event is then numbered within its session:

customer_journey["event_number"] = (
    customer_journey
    .groupby("SessionID")
    .cumcount()
    + 1
)

A session such as:

Session 42

12:01   home
12:03   product_page
12:07   cart
12:10   checkout

therefore becomes:

event_number   page
1              home
2              product_page
3              cart
4              checkout

This makes it possible to construct the first event, first two events, first three events, and so on without using anything that occurred later in the session.

The important change in this experiment is that the prediction is now conditioned on an ordered process history rather than only on a static record or a set of aggregate interaction features.

Sequential Journey Results

Ranking performance did not improve meaningfully as more of the customer journey became visible.

The pre-confirmation prefix models produced:

AUC⁡(1)≈0.457\operatorname{AUC}(1) \approx 0.457 AUC⁡(2)≈0.464\operatorname{AUC}(2) \approx 0.464 AUC⁡(3)≈0.459\operatorname{AUC}(3) \approx 0.459 AUC⁡(4)≈0.450\operatorname{AUC}(4) \approx 0.450

All four results are below the random-ranking benchmark:

AUC⁡=0.5\operatorname{AUC} = 0.5

So this experiment did not support the idea that these customer-journey prefixes recover a stable ranking of eventual conversion.

It did, however, expose an important interpretation problem for any future sequence experiment.

Consider these two partial journeys:

Session A

home
→ product_page
→ cart
Session B

home
→ product_page
→ cart

Both sessions have the same current state: cart.

If a model distinguishes sessions mainly because one has reached cart while another remains on home, then it has learned funnel position. That does not show that the path taken to reach the current state contains additional information.

The distinction can be written as:

P(Yi=1∣Si)P(Y_i = 1 \mid S_i)

versus:

P(Yi=1∣Si,Hi)P(Y_i = 1 \mid S_i, H_i)

where:

  • SiS_i is the current state of session ii
  • HiH_i is the history that led to that state
  • Yi=1Y_i = 1 means the session eventually converts

The first expression asks:

What is the probability of conversion given where the session is now?

The second asks:

What is the probability of conversion given both where the session is now and how it got there?

That is a much stronger test.

Current State as a Baseline

Suppose one customer has reached checkout after moving directly through the funnel:

home
→ product_page
→ cart
→ checkout

Another customer is also currently at checkout, but arrived there after repeated navigation between earlier states:

home
→ product_page
→ home
→ product_page
→ cart
→ product_page
→ cart
→ checkout

A conventional CRM-style snapshot could represent both sessions as:

current_state = checkout

The current-state model therefore receives the same main piece of information for both.

A history-aware representation can preserve the different trajectories.

The relevant comparison becomes:

Performance⁡(Si,Hi)>Performance⁡(Si)\operatorname{Performance}(S_i, H_i) > \operatorname{Performance}(S_i)

If adding HiH_i improves prediction while SiS_i is held constant, then the trajectory itself contains information that is lost in the current-state representation.

This became the basis for the next experiments.

Put simply, the current state is like a summary column, while the history is the set of steps that produced that summary. The next experiments ask whether that summary column has thrown away useful information.

State Transitions

To isolate the effect of history, the next experiment uses an explicit state-transition process.

Instead of predicting only whether a deal eventually closes, the process is represented as a sequence of states:

Si,0→Si,1→Si,2→⋯→Si,tS_{i,0} \rightarrow S_{i,1} \rightarrow S_{i,2} \rightarrow \cdots \rightarrow S_{i,t}

Here, Si,tS_{i,t} is the state of process ii at time tt.

A simplified sales process might look like:

New
→ Contacted
→ Qualified
→ Proposal
→ Won

but another opportunity might follow:

New
→ Contacted
→ Qualified
→ Contacted
→ Qualified
→ Proposal
→ Lost

At some point, both opportunities can occupy the same state even though their histories are different.

That gives us a cleaner question:

Among opportunities that are currently in the same state, does their previous transition history help predict what happens next?

Formally, the baseline uses only the current state:

P(Si,t+1∣Si,t)P(S_{i,t+1} \mid S_{i,t})

The history-aware model uses the current state and the preceding trajectory:

P(Si,t+1∣Si,t,Si,t−1,…,Si,0)P( S_{i,t+1} \mid S_{i,t}, S_{i,t-1}, \ldots, S_{i,0} )

The first expression asks for the probability of the next state given only the current state.

The second asks whether knowing the path taken to reach that state changes the probability of what happens next.

This removes much of the ambiguity from the customer-journey experiment. The comparison is no longer between customers at obviously different points in a funnel. It compares cases that occupy the same current state and asks whether their histories still matter.

That is why the later experiments are stricter than ordinary funnel-stage prediction. They compare like with like: same current state, different prior path.

Within-State Transition History

The previous experiment showed that later process states are easier to rank, but that does not tell us whether the history leading to a state matters.

To test that directly, I built a synthetic state-transition process where opportunities move through a small sales funnel:

S={new,engaged,qualified,proposal,won,lost}\mathcal{S} = \{ \text{new}, \text{engaged}, \text{qualified}, \text{proposal}, \text{won}, \text{lost} \}

Each opportunity starts in new and moves through the process one step at a time.

At any non-terminal state, it can:

  • advance
  • remain in the same state
  • regress
  • become lost

For example:

new
→ engaged
→ qualified
→ proposal
→ won

is one possible trajectory, while:

new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal
→ lost

is another.

A Hidden Propensity

Each synthetic opportunity is assigned an unobserved value:

AiA_i

which affects how likely it is to move forward, stall, regress, or become lost.

Higher values of AiA_i make forward movement more likely and loss less likely.

The transition process can be written as:

P(Si,t+1∣Si,t,Ai)P \left( S_{i,t+1} \mid S_{i,t}, A_i \right)

where:

  • Si,tS_{i,t} is the current state of opportunity ii
  • Si,t+1S_{i,t+1} is its next state
  • AiA_i is the hidden propensity assigned to that opportunity

The expression can be read as:

the probability of the opportunity's next state, given its current state and its unobserved propensity.

The model never receives AiA_i. It only observes what happened to the opportunity.

Generating Transition Probabilities

Each possible transition is first assigned a score.

For example:

zadvance=αs+βaAiz_{\text{advance}} = \alpha_s + \beta_a A_i zstay=0z_{\text{stay}} = 0 zregress=γs−βrAiz_{\text{regress}} = \gamma_s - \beta_r A_i zlost=λs−βlAiz_{\text{lost}} = \lambda_s - \beta_l A_i

The coefficients control how strongly the hidden propensity affects each type of transition.

These scores are not probabilities yet. They can take any real value.

To convert them into probabilities that add up to 11, the experiment uses the softmax function:

P(Ti=j)=exp⁡(zij)∑m∈T(Si,t)exp⁡(zim)P(T_i = j) = \frac{ \exp(z_{ij}) }{ \sum_{m \in \mathcal{T}(S_{i,t})} \exp(z_{im}) }

Here, T(Si,t)\mathcal{T}(S_{i,t}) is the set of transitions that are available from the current state.

For example, an opportunity in qualified might be able to:

advance → proposal
stay    → qualified
regress → engaged
exit    → lost

Softmax converts the four transition scores into four probabilities whose total is:

11

One of those transitions is then sampled, producing the next observed state.

This process repeats until the opportunity reaches either won or lost.

Why Use Synthetic Data Again?

The question is no longer whether a general latent deal score exists.

The synthetic process gives us a controlled environment where we know that past transitions contain information about an unobserved variable.

That lets us test a specific representation question:

If two opportunities are currently in the same state, can their previous transitions help distinguish them?

Suppose two opportunities are both currently qualified.

The first followed:

new
→ engaged
→ qualified

The second followed:

new
→ engaged
→ qualified
→ engaged
→ qualified

A current-state representation gives both of them:

current_state = qualified

A history-aware representation can distinguish the two trajectories.

The comparison is therefore:

P(Yi=1∣Si,t)P(Y_i = 1 \mid S_{i,t})

against:

P(Yi=1∣Si,t,Hi,t)P(Y_i = 1 \mid S_{i,t}, H_{i,t})

where Hi,tH_{i,t} contains the transitions observed before time tt.

If the history-aware model performs better among opportunities occupying the same current state, then current state is losing information about the process that produced it.

That is the specific property being tested here.

Within-State Ranking Results

The history-aware model was evaluated separately within the engaged, qualified, and proposal states.

The results were:

AUC⁡engaged≈0.647\operatorname{AUC}_{\text{engaged}} \approx 0.647 AUC⁡qualified≈0.630\operatorname{AUC}_{\text{qualified}} \approx 0.630 AUC⁡proposal≈0.666\operatorname{AUC}_{\text{proposal}} \approx 0.666

All three are above the random-ranking benchmark:

AUC⁡=0.5\operatorname{AUC} = 0.5

Because each model is evaluated only among opportunities that occupy the same current state, differences in pipeline stage cannot explain the ranking.

For example, every opportunity in the proposal evaluation set is already at proposal. The model therefore cannot perform well simply by learning that proposal is generally closer to won than engaged.

The information available to the model comes from the history that produced that state, including features such as:

  • number of forward transitions
  • number of regressions
  • number of repeated states
  • total number of transitions
  • net progress
  • regression rate
  • most recent transition

So, within the synthetic process, two opportunities can both be at:

proposal

while their histories imply different probabilities of eventually reaching won.

For example:

Opportunity A

new
→ engaged
→ qualified
→ proposal

and:

Opportunity B

new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal

have the same current state but different trajectories.

The result shows that the trajectory contains information that is lost when both opportunities are represented only as:

current_state = proposal

Limitation

This result still comes from a synthetic process.

The hidden propensity AiA_i was deliberately constructed to affect transition probabilities. An opportunity with a higher AiA_i is more likely to advance and less likely to regress or become lost.

Transition history is therefore expected to contain information about AiA_i.

The result does not establish that real sales or business processes behave this way. It shows that if an unobserved factor affects how entities move through a process, their observed transition history can preserve information about that factor even after their current state is known.

That gives us a real-data question:

among cases observed at the same point in a real business process, does the history leading to that point contain useful information about what eventually happens?

Testing Transition History on Real Process Data

The synthetic state-transition experiment shows that history can matter when the process is explicitly constructed so that previous transitions contain information about an unobserved variable.

That still does not tell us whether the same thing happens in real business processes.

The next experiment uses the BPI Challenge 2017 event log, which contains event-level histories for real loan applications. Unlike a normal CRM snapshot, the dataset records the sequence of activities associated with each application over time.

A simplified application history might look like:

A_Create Application
→ A_Submitted
→ A_Concept
→ A_Accepted
→ O_Create Offer
→ O_Sent
→ A_Pending

Another application may reach some of the same intermediate states through a different sequence.

That makes it possible to repeat the within-state test using real process history.

Defining the Outcome

Each application eventually reaches one of three terminal states:

A_Pending
A_Denied
A_Cancelled

For the binary experiment, successful completion is defined as:

Yi={1if application i ends in A_Pending0if application i ends in A_Denied or A_CancelledY_i = \begin{cases} 1 & \text{if application } i \text{ ends in A\_Pending} \\ 0 & \text{if application } i \text{ ends in A\_Denied or A\_Cancelled} \end{cases}

The target is determined only from the final terminal event.

Intermediate activities such as A_Accepted are not treated as successful outcomes because the application can still fail later in the process.

This distinction between intermediate events and final labels is another form of leakage control. The model is not allowed to treat a promising middle state as if it were already the final answer.

Comparing Applications at the Same State

The key requirement is that the applications being compared must occupy the same observable process state.

Suppose two applications have both reached:

O_Sent

At that point, their current state is identical.

For a checkpoint state ss, define:

Cs={i:application i reaches state s}\mathcal{C}_s = \{ i : \text{application } i \text{ reaches state } s \}

Cs\mathcal{C}_s is simply the set of applications that reached the same checkpoint.

For every application in that set, the model is allowed to use only the events that occurred before or at the first occurrence of that checkpoint.

Anything that happens later is hidden.

This gives each application a history:

Hi,t=(Si,0,Si,1,…,Si,t)H_{i,t} = \left( S_{i,0}, S_{i,1}, \ldots, S_{i,t} \right)

while holding the final observed state fixed:

Si,t=sS_{i,t} = s

The question is then:

P(Yi=1∣Si,t=s,Hi,t)P(Y_i = 1 \mid S_{i,t}=s, H_{i,t})

Can the history Hi,tH_{i,t} help rank applications that are all currently at the same state ss?

If it can, then the current state alone is not a sufficient description of the process.

Why the First Occurrence Matters

Some applications can return to the same activity more than once.

For example:

A_Submitted
→ A_Concept
→ A_Accepted
→ A_Concept
→ A_Accepted

If we simply chose an arbitrary occurrence of A_Accepted, some applications would have much more future information available than others.

The experiment therefore uses the first occurrence of the chosen checkpoint.

Everything after that point is excluded.

This gives every application the same prediction condition:

what could have been known the first time this application reached this state?

Temporal Validation

The data is again split chronologically.

Applications that started earlier are used for training, while later applications are reserved for testing.

This matters because a random split could allow the model to learn process patterns from future cases and use them to predict earlier ones.

The real question is whether historical process data can help rank applications that occur later in time.

Within-State AUC

For each checkpoint state ss, ranking performance is measured separately:

AUC⁡s\operatorname{AUC}_s

Because every application in that evaluation set already occupies the same current state, the state itself cannot explain differences in ranking.

A value of:

AUC⁡s=0.5\operatorname{AUC}_s = 0.5

would mean the history provides no useful ordering within that state.

A value above:

0.50.5

would indicate that the path leading to the checkpoint contains information about the eventual terminal outcome.

That is a much stronger test than asking whether applications further along in the process are more likely to succeed.

Real Within-State Results

The first history-only models used aggregate process features such as event counts, elapsed time, repeated activities, offer activity, and workflow activity.

The A_Submitted checkpoint was excluded before modeling because almost every application reached it through the same observed history. There was essentially no variation to rank on.

The remaining checkpoints were:

  • A_Validating
  • A_Incomplete

The history-only models produced:

AUC⁡A_Validating≈0.581\operatorname{AUC}_{A\_Validating} \approx 0.581

and:

AUC⁡A_Incomplete≈0.589\operatorname{AUC}_{A\_Incomplete} \approx 0.589

Both are above the random-ranking benchmark:

AUC⁡=0.5\operatorname{AUC} = 0.5

Because each model is evaluated only among applications at the same checkpoint, the result cannot be explained by one application simply being further through the process than another.

The difference comes from what happened before the checkpoint.

The effect is still fairly weak. An AUC around 0.580.58 does not support accurate within-state ranking, but it does suggest that some information about the eventual outcome survives in the prior process history.

Preserving More of the Sequence

The first real-data model still compresses history into aggregate counts.

For example, these histories:

A → B → C → B → C

and:

A → B → B → C → C

can produce similar event counts despite having different transition structures.

The next version therefore added features describing adjacent transitions:

(ei1→ei2),(ei2→ei3),…,(ei,n−1→ein)(e_{i1} \rightarrow e_{i2}), (e_{i2} \rightarrow e_{i3}), \ldots, (e_{i,n-1} \rightarrow e_{in})

Instead of recording only how often an activity occurred, the model could also record transitions such as:

A_Validating → A_Incomplete
A_Incomplete → W_Call incomplete files
O_Create Offer → O_Sent

The sequence-aware model produced:

AUC⁡A_Validating≈0.592\operatorname{AUC}_{A\_Validating} \approx 0.592

and:

AUC⁡A_Incomplete≈0.568\operatorname{AUC}_{A\_Incomplete} \approx 0.568

Compared with the aggregate-history model:

0.581→0.5920.581 \rightarrow 0.592

for A_Validating, while:

0.589→0.5680.589 \rightarrow 0.568

for A_Incomplete.

Preserving adjacent transition structure therefore improved one checkpoint slightly and made the other worse.

There is some recoverable information in the process history, but simply adding more sequence detail does not produce a consistently stronger ranking model.

What This Actually Supports

The original conjecture was much broader: that opportunity outcomes might be recoverable through a generalized ranking function representing some latent transaction propensity.

These experiments do not provide strong evidence for that claim.

The real-data results are much narrower. Applications observed at the same process state can still differ in ways that are partially recoverable from the history that brought them there.

A useful way to state that is:

Si,t=g(Hi,t)S_{i,t} = g(H_{i,t})

where the current state Si,tS_{i,t} is produced by some preceding history Hi,tH_{i,t}.

Representing only Si,tS_{i,t} therefore compresses that history into a much smaller description of the process.

Different histories can map to the same current state:

Ha≠HbH_a \neq H_b

while:

g(Ha)=g(Hb)g(H_a) = g(H_b)

Some information that distinguishes HaH_a from HbH_b is necessarily absent from the current-state representation.

The empirical question is whether any of that discarded information matters for the decision being made.

In the BPI experiment, at least some of it does.

The effect is modest, but the result is enough to motivate a different prediction target: rather than trying to predict a distant terminal outcome such as won, lost, or successful completion, can the current process history predict what happens next?

Next-State Prediction

The terminal-outcome experiments produced only modest within-state ranking performance. The next experiment changes the target.

Instead of asking whether an application will eventually succeed, the model predicts the next observable process event.

For case ii at time tt, let:

Si,tS_{i,t}

represent the current state,

Hi,tH_{i,t}

the process history observed so far, and:

Gi,tG_{i,t}

the relational context surrounding the case.

The target is:

Si,t+1S_{i,t+1}

so the full prediction problem becomes:

P(Si,t+1∣Si,t,Hi,t,Gi,t)P( S_{i,t+1} \mid S_{i,t}, H_{i,t}, G_{i,t} )

This can be read as:

the probability of the next process state, given the current state, the history so far, and the surrounding relational context.

The notation is compact, but the idea is ordinary supervised learning: each row contains what was known at time tt, and the label is what happened at time t+1t+1. Appendix A5 translates this into a simple table-style example.

Constructing Prediction Snapshots

The BPI event log is first ordered by application and timestamp:

bpi = (
    bpi
    .sort_values(
        ["case:concept:name", "time:timestamp"]
    )
    .reset_index(drop=True)
)

Each event is then assigned its position within the application:

bpi["event_position"] = (
    bpi
    .groupby("case:concept:name")
    .cumcount()
)

The next event is created by shifting the activity column one row backward within each application:

bpi["next_activity"] = (
    bpi
    .groupby("case:concept:name")["concept:name"]
    .shift(-1)
)

This turns an event sequence such as:

A_Submitted
→ A_Concept
→ A_Accepted
→ O_Create Offer

into supervised examples:

current: A_Submitted
target:  A_Concept
current: A_Concept
target:  A_Accepted
current: A_Accepted
target:  O_Create Offer

The last event in each application is removed because there is no following event to predict:

event_snapshots = bpi[
    bpi["next_activity"].notna()
].copy()

Temporal Train/Test Split

The split is performed at the application level using case start time.

First, the start time of each application is calculated:

case_start_times = (
    bpi
    .groupby("case:concept:name")["time:timestamp"]
    .min()
)

A chronological cutoff is then chosen:

cutoff = case_start_times.quantile(0.75)

Earlier applications form the training set:

train_snapshots = event_snapshots[
    event_snapshots["case_start_time"] < cutoff
].copy()

Later applications form the test set:

test_snapshots = event_snapshots[
    event_snapshots["case_start_time"] >= cutoff
].copy()

This preserves the same basic rule used throughout the experiments: the model should learn from the past and be evaluated on later cases.

Appendix A7 gives the general version of this rule, and Appendix B1, Appendix B4, and Appendix B6 explain why the reported next-event metrics are accuracy, macro F1, and weighted F1 rather than only AUC.

State-Only Baseline

The first baseline uses only the current activity.

For every state ss, the training data is used to find the next event that followed that state most often:

S^t+1(s)=arg⁡max⁡jP(St+1=j∣St=s)\hat{S}_{t+1}(s) = \arg\max_j P(S_{t+1}=j \mid S_t=s)

The arg⁡max⁡\arg\max operator means:

choose the value of jj that gives the largest probability.

In code:

state_next_mode = (
    train_snapshots
    .groupby("concept:name")["next_activity"]
    .agg(
        lambda x: x.value_counts().idxmax()
    )
)

If A_Concept was most often followed by A_Accepted in the historical training data, then the baseline predicts:

A_Concept → A_Accepted

whenever it encounters A_Concept.

Those predictions are applied to the test set:

test_snapshots["state_only_prediction"] = (
    test_snapshots["concept:name"]
    .map(state_next_mode)
)

The resulting temporal accuracy was:

Accuracy⁡state-only≈0.666\operatorname{Accuracy}_{\text{state-only}} \approx 0.666

with complete coverage of the test set.

So current state alone correctly predicts the next event about two-thirds of the time.

That is already much stronger than the earlier attempts to predict a distant terminal outcome.

Adding Process History

The next model includes compact information about the path observed so far.

The history representation includes:

  • event position
  • elapsed time since application creation
  • previous activity
  • number of unique activities observed
  • number of repeated activities
  • cumulative application-event count
  • cumulative offer-event count
  • cumulative workflow-event count

The prediction changes from:

P(St+1∣St)P(S_{t+1} \mid S_t)

to:

P(St+1∣St,Ht)P(S_{t+1} \mid S_t, H_t)

A simplified feature set might look like:

current_activity        = A_Validating
previous_activity       = A_Incomplete
event_position          = 12
elapsed_hours           = 47.3
unique_activity_count   = 8
repeated_activity_count = 4
offer_event_count       = 2
workflow_event_count    = 5

Categorical features such as current_activity and previous_activity are one-hot encoded. Numerical features are standardized. Appendix A6 shows why both transformations are needed before a standard classifier can read the data.

The model is built with scikit-learn:

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.linear_model import SGDClassifier

The preprocessing step can be defined as:

preprocessor = ColumnTransformer(
    transformers=[
        (
            "categorical",
            OneHotEncoder(handle_unknown="ignore"),
            categorical_features,
        ),
        (
            "numeric",
            StandardScaler(),
            numeric_features,
        ),
    ]
)

The classifier and preprocessing are then combined into a single pipeline:

model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        (
            "classifier",
            SGDClassifier(
                loss="log_loss",
                max_iter=2000,
                random_state=42,
            ),
        ),
    ]
)

The important point is that preprocessing is learned only from the training data.

The model is then fit with:

model.fit(
    X_train,
    y_train
)

and next-event predictions are generated for the temporal holdout:

history_prediction = model.predict(
    X_test
)

The resulting accuracy was:

Accuracy⁡state+history≈0.781\operatorname{Accuracy}_{\text{state+history}} \approx 0.781

Compared with the state-only baseline:

0.666→0.7810.666 \rightarrow 0.781

which is an absolute improvement of:

0.1150.115

or about 11.5 percentage points.

The history of the process is therefore much more useful for predicting the next event than it was for predicting a distant final outcome.

Relational Context Ablation

The next step was to test which additional forms of context actually mattered.

Rather than adding every available field at once, the experiment uses ablation.

Ablation means training several otherwise similar models while adding or removing specific groups of features. If performance changes materially when one group is introduced, that group is contributing useful predictive information.

The main representations were:

State + History
State + History + Case
State + History + Resource
State + History + Case + Resource

The models were evaluated with:

  • accuracy
  • macro F1
  • weighted F1

These are classification metrics rather than ranking metrics because the target is now a specific next event, not a binary win/loss outcome. Appendix B1 explains the difference.

The results were:

Representation Accuracy Macro F1 Weighted F1
State + History 0.781 0.604 0.751
State + History + Case 0.786 0.613 0.757
State + History + Resource 0.838 0.679 0.820
State + History + Case + Resource 0.840 0.683 0.823

Adding case-level information produced only a small improvement.

Adding resource identity produced a much larger one:

0.781→0.8380.781 \rightarrow 0.838

What Does "Resource" Mean?

In the BPI process log, events can be associated with the resource responsible for that activity.

That might represent a user, employee, team member, or operational handler.

So a row can contain information conceptually like:

case      = Application_4821
activity  = A_Validating
resource  = User_37

When resource identity is added to the model, it can learn that different resources are associated with different transition patterns.

That does not establish that the resource causes the transition.

Resource identity can also capture:

  • specialization
  • role
  • assignment rules
  • case complexity
  • organizational structure
  • operational routing
  • workload patterns

The ablation establishes only that this information helps predict what happens next. Metric definitions for accuracy, macro F1, and weighted F1 are included in Appendix B4, Appendix B5, and Appendix B6.

Graph-Derived Context

Resource identity is relational information, but encoding a resource as a category still treats it as a label.

For example:

resource = User_37

does not tell the model how User_37 is connected to other applications or resources.

The event log can also be represented as a graph.

A simplified local structure might look like:

Application A
├── handled by → Resource 12
├── handled by → Resource 41
└── contains   → Event ...

Application B
├── handled by → Resource 12
└── contains   → Event ...

Because Resource 12 appears in both applications, the two cases are connected through a shared node.

Graph-derived features summarize parts of this local structure and add them to the existing state, history, and resource representation.

At this point, the feature vector is no longer just a set of fields from one row. It also includes simple summaries of the case's local graph context.

The final comparison was:

Representation Accuracy Macro F1 Weighted F1
State only 0.666 0.483 0.609
State + History 0.781 0.604 0.751
State + History + Resource 0.838 0.679 0.820
State + History + Resource + Graph 0.846 0.692 0.829

The graph-derived features produced a smaller additional improvement:

0.838→0.8460.838 \rightarrow 0.846

in accuracy, while macro F1 increased:

0.679→0.6920.679 \rightarrow 0.692

The full progression was:

0.666→0.781→0.838→0.8460.666 \rightarrow 0.781 \rightarrow 0.838 \rightarrow 0.846

The largest improvement came from process history.

Resource identity added another substantial increase.

Graph-derived relational context added a smaller but consistent improvement across the reported metrics.

What This Suggests for CRM Architecture

The original appetite conjecture did not survive the experiments in a useful general form.

Static CRM data performed poorly on future opportunity ranking. Adding interaction summaries helped only modestly. Sequential history became more useful once the target moved closer to the process itself, and the strongest result came from predicting the next transition using current state, process history, resource identity, and local relational context.

The progression in the final experiment was:

0.666→0.781→0.838→0.8460.666 \rightarrow 0.781 \rightarrow 0.838 \rightarrow 0.846

for:

current state
→ state + history
→ state + history + resource
→ state + history + resource + graph context

This does not show that every CRM should use a graph database.

It does suggest that the representation used for prediction matters.

The Relational Representation

A conventional relational CRM might store an opportunity across tables such as:

opportunities
-------------
id
account_id
owner_id
stage
value
created_at
updated_at

with related information in other tables:

contacts
activities
users
accounts
documents
stage_history

There is nothing inherently incapable about this representation. SQL can reconstruct relationships and history through joins.

For example:

SELECT
    o.id,
    o.stage,
    a.type,
    u.id AS owner_id,
    h.previous_stage,
    h.changed_at
FROM opportunities o
LEFT JOIN accounts a
    ON a.id = o.account_id
LEFT JOIN users u
    ON u.id = o.owner_id
LEFT JOIN stage_history h
    ON h.opportunity_id = o.id;

The issue is not whether a relational database can store the information. It can.

The question is what representation should be exposed to a decision model when the decision depends on a connected process unfolding over time.

A single opportunity row might say:

opportunity_id = 1842
stage          = proposal
owner          = user_17
value          = 120000

but the process that produced that row may contain:

created
→ assigned to user_04
→ contacted
→ qualified
→ assigned to user_17
→ proposal created
→ customer replied
→ proposal revised
→ proposal

The current row is one projection of that history.

A Graph Representation

The same process can be represented in terms of entities and their relationships:

Opportunity 1842
├── BELONGS_TO → Account 93
├── OWNED_BY → User 17
├── INVOLVES → Contact 51
├── HAS_EVENT → Call 882
├── HAS_EVENT → Email 901
├── HAS_EVENT → Proposal 44
└── HAS_STATE → Proposal

Events can also connect to each other through time:

Call 882
→ FOLLOWED_BY → Email 901
→ FOLLOWED_BY → Proposal 44
→ FOLLOWED_BY → Revision 46

and resources can connect multiple cases:

User 17
├── HANDLED → Opportunity 1842
├── HANDLED → Opportunity 2201
└── HANDLED → Opportunity 2390

That makes relationships part of the representation itself rather than something reconstructed only when a particular query needs them.

For a prediction at time tt, the available context can be thought of as a local subgraph:

Gi,t=(Vi,t,Ei,t)G_{i,t} = (V_{i,t}, E_{i,t})

where:

  • Vi,tV_{i,t} is the set of relevant entities known at time tt
  • Ei,tE_{i,t} is the set of relationships between those entities

A model can then condition on:

P(Si,t+1∣Si,t,Hi,t,Gi,t)P( S_{i,t+1} \mid S_{i,t}, H_{i,t}, G_{i,t} )

The experiments do not establish that this is always the best representation. They show that information represented by Hi,tH_{i,t} and parts of Gi,tG_{i,t} improved next-state prediction on the process data tested here.

Appendix A2 and Appendix A6 give the preprocessing view of this: graph context still has to be converted into model-readable features before a standard classifier can use it.

Current State Is a Compression

The clearest result from the experiments is that current state can discard information.

Let the complete process history before time tt be:

Hi,t=(ei,0,ei,1,…,ei,t)H_{i,t} = ( e_{i,0}, e_{i,1}, \ldots, e_{i,t} )

and let the CRM derive a current state from that history:

Si,t=g(Hi,t)S_{i,t} = g(H_{i,t})

Many different histories can produce the same state.

For two cases aa and bb:

Ha,t≠Hb,tH_{a,t} \neq H_{b,t}

while:

g(Ha,t)=g(Hb,t)g(H_{a,t}) = g(H_{b,t})

A simple example is:

Case A

new
→ engaged
→ qualified
→ proposal

and:

Case B

new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal

Both can be stored as:

stage = proposal

but the histories are different.

Whether that difference matters depends on the task.

For terminal-outcome ranking in the real data, the additional information was weak.

For next-state prediction, it mattered considerably more.

That distinction is important. There is no reason to assume that richer process state will improve every prediction problem equally.

A Stateful CRM

The architecture I find more interesting after these experiments is a CRM where the current record remains useful, but it is treated as one view over an underlying process history.

Conceptually:

entities
    +
events
    +
relationships
    +
time
    ↓
current process state
    ↓
prediction / decision

An event might look like:

{
  "type": "proposal_sent",
  "opportunity_id": "1842",
  "actor_id": "user_17",
  "contact_id": "contact_51",
  "timestamp": "2026-08-31T15:42:00Z"
}

Another event could record a state transition:

{
  "type": "stage_changed",
  "opportunity_id": "1842",
  "from": "qualified",
  "to": "proposal",
  "actor_id": "user_17",
  "timestamp": "2026-08-31T15:44:00Z"
}

The current opportunity record can still be derived and queried normally:

stage = proposal
owner = user_17

but the underlying history remains available when a decision depends on how that state was reached.

This also avoids requiring every downstream model to treat the CRM row as a complete description of the business process.

What Survived

I started these experiments looking for a generalized ranking function for deal appetite.

The synthetic version was recoverable because the latent structure had been placed into the simulation. Real CRM data did not provide evidence for a comparably useful scalar, and increasingly detailed histories did not produce a strong general model of terminal outcomes.

The more consistent result concerned process state.

Static CRM fields were weak for future deal ranking. Real transition history retained modest information within the same process state. When the target was changed to the next observable transition, history produced a much larger improvement, resource identity added more, and graph-derived relational context added a smaller additional gain.

The experiments therefore give me a narrower hypothesis to work with:

Business state may be better represented for decision-making as a history of connected entities and events than as the current values of those entities alone.

That is a claim about representation, not about Neo4j specifically.

A relational database can store the same underlying information. A graph may simply provide a more direct model when the questions being asked depend on paths, shared entities, transitions, and local context.

The remaining question is whether that advantage survives outside this dataset and whether it is large enough to justify the additional complexity of a graph-native representation.

Hypothesis Revisited

The original hypotheses were:

H1H_1

There exists a recoverable ordering function over observable CRM variables:

Ai=f(xi)+εiA_i = f(x_i) + \varepsilon_i

such that f(xi)f(x_i) preserves useful relative ordering even when εi\varepsilon_i is unobserved.

H0H_0

Observable CRM variables do not recover a stable relative ordering of opportunities beyond noise and unobserved variation.

The experiments do not provide enough evidence to support H1H_1 in this general form.

The synthetic experiment recovered the ranking structure that had been explicitly built into the data-generating process. That established that the estimation procedure could recover the systematic component under the assumptions of the simulation, but it did not establish that a comparable one-dimensional ordering function exists in real CRM data.

On real CRM data, static opportunity features produced weak out-of-sample ranking of terminal outcomes. Interaction history and process-history features improved performance in some experiments, but the gains were not strong or consistent enough to establish a stable generalized ordering of opportunities.

I therefore would not reject H0H_0 for the original conjecture.

The later experiments support a narrower result.

Process history can retain information that is absent from the current-state representation, and that information becomes substantially more useful when the target is the next observable process transition rather than a distant terminal outcome.

In the final experiment, next-event accuracy increased from:

0.6660.666

using current state alone, to:

0.7810.781

after adding process history, then to:

0.8380.838

after adding resource identity, and finally to:

0.8460.846

after adding graph-derived relational context.

Those results do not rescue the original H1H_1. They point to a different hypothesis about how business processes should be represented for prediction and decision-making.

Conclusion

The deal-appetite conjecture began as a search for a generalized opportunity-ranking function. The synthetic experiment showed that such a function can be recovered when the recoverable structure is deliberately built into the data-generating process. The real-data experiments did not show that an analogous scalar exists in ordinary CRM records.

The more durable finding is about representation.

When the target is a distant terminal outcome, richer history is only weakly useful in the datasets tested here. When the target is the next observable process transition, history and relational context matter much more.

That suggests a narrower and more useful direction:

Business prediction should often be framed around the stateful process that produces the current CRM record, not only around the current record itself.

A graph database is not required for that claim to be true. But graph-shaped representations make the relevant objects easier to name directly: events, actors, accounts, contacts, offers, states, transitions, and the paths between them.

The next question is whether richer graph-native representations can improve decision support in real CRM systems enough to justify their added complexity.

Appendices

These appendices are numbered by topic. Appendix A covers notation, feature construction, and preprocessing. Appendix B covers evaluation metrics. Appendix C covers data and reproducibility.

Appendix A1: Libraries and Notation

This appendix explains the notation and library mechanics used throughout the experiments. It is meant to make the article readable even if the reader has only seen introductory statistics or discrete math.

The experiments were written in Python and run in Jupyter.

The main libraries were:

  • pandas for loading, joining, filtering, grouping, and transforming CRM and event-log data
  • NumPy for random-number generation and vectorized numerical operations in the synthetic experiments
  • scikit-learn for preprocessing, model fitting, and evaluation
  • SciPy for Spearman rank correlation
  • Matplotlib for plots and visual inspection of results

Notation guide

The notation is compact because the same structure appears many times.

Symbol Plain meaning
ii the index of one opportunity, application, case, or event snapshot
jj another index, usually used when comparing two examples
nn the number of examples
xx one input value
xi\mathbf{x}_i the full feature vector for example ii
XX the full feature table or matrix
yiy_i the observed label for example ii
p^i\hat{p}_i the model's estimated probability for example ii
y^i\hat{y}_i the model's predicted class for example ii
SiS_i systematic appetite in the synthetic experiment
AiA_i total latent appetite in the synthetic experiment
ϵi\epsilon_i unobserved variation for example ii
ρs\rho_s Spearman rank correlation

A subscript tells us which example we are talking about.

For example:

x7\mathbf{x}_7

means "the feature vector for example 7."

A superscript in this article is usually a label, not a power.

For example:

xi(s)\mathbf{x}_i^{(s)}

means the static CRM feature vector for opportunity ii.

It does not mean xi\mathbf{x}_i multiplied by itself.

The hat symbol means estimated by the model.

For example:

pip_i

means the true probability in the synthetic experiment, while:

p^i\hat{p}_i

means the model's estimated probability.

The vertical bar means "given" or "conditional on."

For example:

P(Yi=1∣xi)P(Y_i = 1 \mid \mathbf{x}_i)

means:

the probability that example ii has outcome 11, given the features in xi\mathbf{x}_i.

The membership symbol means "is an element of."

For example:

σ∈[0,5]\sigma \in [0,5]

means:

σ\sigma is chosen from the interval between 00 and 55.

In the noise-sensitivity experiment, this was implemented as a finite grid:

noise_levels = np.arange(0, 5.25, 0.25)

So the computer did not test every real number between 00 and 55. It tested:

0.00, 0.25, 0.50, ..., 5.00

The normal-distribution notation:

xi∼N(0,1)x_i \sim \mathcal{N}(0,1)

means:

draw xix_i from a normal distribution with mean 00 and standard deviation 11.

Appendix A2: Feature Vectors

A model cannot read a CRM row the way a person reads a spreadsheet. It needs a list of values.

For one opportunity ii, that list is called a feature vector:

xi\mathbf{x}_i

The bold x\mathbf{x} means that the input contains several values, not just one value.

Suppose one opportunity has:

product          = GTX Pro
sector           = Technology
sales_price      = 82,000
employees        = 220
is_subsidiary    = 0
engage_month     = 3

Before a standard model can use that row, text values must be converted into numbers. After preprocessing, the model might receive this vector:

feature name          value
-------------------   -----
product_GTX_Basic       0
product_GTX_Pro         1
product_MG_Special      0
sector_Finance          0
sector_Technology       1
scaled_sales_price      0.42
scaled_employees       -0.18
is_subsidiary           0
engage_month            3

The same thing can be written as a mathematical vector:

xi=[010010.42−0.1803]\mathbf{x}_i = \begin{bmatrix} 0 \\ 1 \\ 0 \\ 0 \\ 1 \\ 0.42 \\ -0.18 \\ 0 \\ 3 \end{bmatrix}

Each row of the vector is one input feature.

When there are many opportunities, the model receives a table of feature vectors:

X=[−−−x1−−−−−−x2−−−⋯−−−xn−−−]X = \begin{bmatrix} --- \mathbf{x}_1 --- \\ --- \mathbf{x}_2 --- \\ \cdots \\ --- \mathbf{x}_n --- \end{bmatrix}

Each row is one example. Each column is one feature.

A small feature table might look like this:

opportunity product_GTX_Pro sector_Technology scaled_sales_price won
1 1 1 0.42 1
2 0 1 -0.60 0
3 1 0 1.10 1

The model is trained on the feature columns:

XX

and the label column:

yy

For a binary sales model:

y_i = 1  if the deal was won
y_i = 0  if the deal was lost

So the model learns a relationship of the form:

xi→yi\mathbf{x}_i \rightarrow y_i

or, for probability prediction:

xi→p^i\mathbf{x}_i \rightarrow \hat{p}_i

where p^i\hat{p}_i is the model's estimated probability.

Appendix A3: Feature Groups

The article uses superscripts to label different groups of inputs.

For example:

xi(s)\mathbf{x}_i^{(s)}

means static CRM features, such as product, sector, salesperson, and deal value.

xi(b)\mathbf{x}_i^{(b)}

means behavioral or interaction-derived features, such as interaction count, reply timing, or channel mix.

The combined model can therefore be read as:

P(Yi=1∣xi(s),xi(b))P \left( Y_i = 1 \mid \mathbf{x}_i^{(s)}, \mathbf{x}_i^{(b)} \right)

which means:

the probability that opportunity ii is won, given its static CRM features and behavioral features.

Appendix A4: Logistic Regression

In a binary model, logistic regression first converts the feature vector into a single score:

zi=β0+β1xi1+β2xi2+⋯+βkxikz_i = \beta_0 + \beta_1x_{i1} + \beta_2x_{i2} + \cdots + \beta_kx_{ik}

Here:

  • xi1x_{i1} is the first feature for example ii
  • xi2x_{i2} is the second feature for example ii
  • kk is the number of features
  • β0\beta_0 is the intercept
  • β1,…,βk\beta_1,\ldots,\beta_k are learned coefficients

A tiny worked example:

features:
    product_GTX_Pro      = 1
    sector_Technology    = 1
    scaled_sales_price   = 0.42

learned coefficients:
    intercept            = -0.30
    product_GTX_Pro      = 0.50
    sector_Technology    = 0.20
    scaled_sales_price   = 0.80

The model score is:

zi=−0.30+0.50(1)+0.20(1)+0.80(0.42)=0.736z_i = -0.30 + 0.50(1) + 0.20(1) + 0.80(0.42) = 0.736

That score is not yet a probability. Logistic regression applies the logistic function:

p^i=11+e−zi\hat{p}_i = \frac{1}{1 + e^{-z_i}}

So:

p^i=11+e−0.736≈0.676\hat{p}_i = \frac{1}{1 + e^{-0.736}} \approx 0.676

The model would estimate about a 67.6%67.6\% probability of the positive class.

In code, scikit-learn does this after the model has been fit:

predicted_probability = model.predict_proba(X_test)[:, 1]

The [:, 1] means:

take every row, and take column 1 of the probability output.

For a binary classifier, column 0 is usually the probability of class 0, and column 1 is the probability of class 1.

Appendix A5: Histories as Feature Vectors

For process-history models, the raw object is no longer just one CRM row. It is a history:

A_Submitted
-> A_Concept
-> A_Accepted
-> O_Create Offer

To use a standard classifier, that history still has to be converted into features.

For example:

current_activity        = A_Accepted
previous_activity       = A_Concept
event_position          = 3
elapsed_hours           = 12.7
unique_activity_count   = 3
offer_event_count       = 0
workflow_event_count    = 1

Those values form the feature vector for that prediction time.

For next-event prediction, each training row contains information available at time tt, and the label is the event that happened at time t+1t+1.

case event position tt current activity previous activity elapsed hours label: next activity
C-001 3 A_Accepted A_Concept 12.7 O_Create Offer
C-002 5 W_Complete application W_Handle leads 3.4 A_Complete

This is the same supervised-learning structure as before:

known inputs at time t→observed outcome at time t+1\text{known inputs at time } t \rightarrow \text{observed outcome at time } t+1

The important rule is that features must be computed only from information available at or before time tt.

Appendix A6: Categorical and Numerical Preprocessing

For static CRM models, categorical and numerical fields require different preprocessing.

Categorical fields are fields with names or categories, such as:

product = GTX Pro
sector  = Technology
office  = West

They were converted into numeric columns using OneHotEncoder:

from sklearn.preprocessing import OneHotEncoder

For example:

product = GTX Pro

may become:

GTX Basic   GTX Pro   MG Special
    0          1           0

The setting:

OneHotEncoder(handle_unknown="ignore")

means:

if a category appears in the test set that was not seen during training, do not crash; encode it as all zeros for that feature group.

Numerical fields are already numbers, such as revenue, employee count, elapsed time, or event count.

They were standardized using StandardScaler:

from sklearn.preprocessing import StandardScaler

Standardization converts a raw value xx into:

z=x−μσz = \frac{x - \mu}{\sigma}

where μ\mu is the mean of that feature in the training data and σ\sigma is its standard deviation.

For example, if the average deal value in the training set is 50,00050{,}000 and the standard deviation is 20,00020{,}000, then a deal worth 70,00070{,}000 becomes:

z=70,000−50,00020,000=1z = \frac{70{,}000 - 50{,}000}{20{,}000} = 1

That means the deal is one training-set standard deviation above the training-set average.

A deal worth 30,00030{,}000 would become:

z=30,000−50,00020,000=−1z = \frac{30{,}000 - 50{,}000}{20{,}000} = -1

That means the deal is one training-set standard deviation below the training-set average.

The two preprocessing steps were combined using ColumnTransformer:

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler

preprocessor = ColumnTransformer(
    transformers=[
        (
            "categorical",
            OneHotEncoder(handle_unknown="ignore"),
            categorical_features,
        ),
        (
            "numeric",
            StandardScaler(),
            numeric_features,
        ),
    ]
)

This says:

apply OneHotEncoder to categorical_features
apply StandardScaler to numeric_features
combine the resulting columns into one model-readable matrix

Appendix A7: Training-Only Preprocessing

Most models used scikit-learn pipelines so that preprocessing was fit only on the training data and then applied unchanged to the temporal holdout.

This matters because preprocessing itself can leak future information.

Suppose the model is evaluated on later opportunities, but the scaler is fit using both earlier and later opportunities. Then the training process has already seen information about the distribution of the future test set.

That gives the model information it would not have had at the time of prediction.

The safer workflow is:

split rows into train and test by time
fit preprocessing on training rows only
transform training rows
train model
transform future test rows using the same fitted preprocessing
evaluate model

The preprocessing step and classifier were wrapped in a Pipeline:

from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression

pipeline = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=2000)),
    ]
)

Training and prediction then followed the standard scikit-learn pattern:

pipeline.fit(X_train, y_train)

predicted_probability = pipeline.predict_proba(X_test)[:, 1]

Inside pipeline.fit, scikit-learn does this:

preprocessor.fit_transform(X_train)
classifier.fit(transformed_X_train, y_train)

Inside pipeline.predict or pipeline.predict_proba, it does this:

preprocessor.transform(X_test)
classifier.predict(...) or classifier.predict_proba(...)

Notice that the test data is transformed, but the preprocessing rules are not refit on the test data.

For the next-event prediction experiment, the target was multiclass rather than binary. Those models used an SGDClassifier with logistic loss:

from sklearn.linear_model import SGDClassifier

classifier = SGDClassifier(
    loss="log_loss",
    max_iter=1000,
    tol=1e-3,
    random_state=42,
)

In this setting, the classifier is not asking "won or lost?" It is asking:

which activity is most likely to happen next?

So the prediction is a class label:

prediction = pipeline.predict(X_test)

and the evaluation uses multiclass metrics.

Appendix B1: Evaluation Tasks

The experiments use two kinds of prediction tasks.

The first kind is binary ranking: rank won deals above lost deals, or successful applications above unsuccessful applications.

The second kind is multiclass classification: predict exactly which event happens next.

Those tasks need different metrics.

Appendix B2: ROC AUC

For binary opportunity-outcome experiments, ranking performance was measured with ROC AUC:

from sklearn.metrics import roc_auc_score

auc = roc_auc_score(
    y_test,
    predicted_probability
)

ROC AUC measures how often the model ranks a randomly selected positive case above a randomly selected negative case.

Here is a small worked example.

deal observed outcome model score
A 1 0.90
B 0 0.80
C 1 0.70
D 0 0.20

There are two positive examples, A and C, and two negative examples, B and D.

That creates four positive-negative pairs:

pair correct ranking? reason
A vs B yes 0.90>0.800.90 > 0.80
A vs D yes 0.90>0.200.90 > 0.20
C vs B no 0.70<0.800.70 < 0.80
C vs D yes 0.70>0.200.70 > 0.20

Three of the four pairs are ordered correctly, so:

AUC⁡=34=0.75\operatorname{AUC} = \frac{3}{4} = 0.75

With ties, AUC gives half credit. The general counting version is:

AUC⁡=#correct pairs+0.5⋅#tied pairs#positive examples⋅#negative examples\operatorname{AUC} = \frac{ \#\text{correct pairs} + 0.5 \cdot \#\text{tied pairs} }{ \#\text{positive examples} \cdot \#\text{negative examples} }

This is why AUC is a ranking metric. It does not require the predicted probability to be perfectly calibrated. It asks whether positive examples tend to receive higher scores than negative examples.

An AUC of:

0.50.5

is equivalent to random ranking, while:

1.01.0

represents perfect ranking.

In the synthetic experiment, AUC was useful because the observed outcome YiY_i was stochastic. Even if a deal had high true probability, it could still fail to close. AUC asks whether the model tends to rank the closed deals higher overall.

Appendix B3: Spearman Rank Correlation

The synthetic experiment also compares the model's predicted ranking against latent quantities that are known only because the data was simulated.

This was measured using Spearman rank correlation:

from scipy.stats import spearmanr

rank_corr, _ = spearmanr(
    predicted_probability,
    test_appetite
)

Spearman correlation first converts values into ranks.

For example:

raw scores:   0.20   0.90   0.40
ranks:          1      3      2

Then it compares the rank ordering of one variable with the rank ordering of another variable.

Here is a worked example with four opportunities.

opportunity model score p^\hat{p} rank of p^\hat{p} latent appetite AA rank of AA rank difference dd
A 0.10 1 5.0 1 0
B 0.85 4 8.0 3 1
C 0.40 2 6.0 2 0
D 0.70 3 10.0 4 -1

Without ties, Spearman correlation can be computed as:

ρs=1−6∑idi2n(n2−1)\rho_s = 1 - \frac{ 6\sum_i d_i^2 }{ n(n^2 - 1) }

In this example:

∑idi2=02+12+02+(−1)2=2\sum_i d_i^2 = 0^2 + 1^2 + 0^2 + (-1)^2 = 2

and:

ρs=1−6(2)4(42−1)=1−1260=0.8\rho_s = 1 - \frac{6(2)}{4(4^2 - 1)} = 1 - \frac{12}{60} = 0.8

A value near:

11

means the two rankings are almost the same.

A value near:

00

means there is no consistent ranking relationship.

A value near:

−1-1

means the rankings are almost opposite.

This is why the synthetic experiment can distinguish:

ρs(p^,A)\rho_s(\hat{p}, A)

from:

ρs(p^,S)\rho_s(\hat{p}, S)

The first compares the model ranking with total latent appetite. The second compares it with the systematic ordering recoverable from observable features.

This distinction matters because:

Ai=Si+ϵiA_i = S_i + \epsilon_i

The model can learn structure in SiS_i because SiS_i is built from observable inputs. It cannot directly recover the realized value of ϵi\epsilon_i because that term is hidden.

Appendix B4: Accuracy

For next-event prediction, the target has many possible classes, so the experiment reports accuracy, macro F1, and weighted F1.

Accuracy measures the proportion of predictions that are exactly correct.

For example, if the model makes 100100 predictions and gets 8484 right:

Accuracy⁡=84100=0.84\operatorname{Accuracy} = \frac{84}{100} = 0.84

In scikit-learn:

from sklearn.metrics import accuracy_score

accuracy = accuracy_score(
    y_test,
    prediction
)

Accuracy is easy to understand, but it can hide poor performance on rare classes.

Suppose a dataset has 100100 next events:

80 are A_Submitted
15 are O_Create Offer
5 are A_Cancelled

A model that mostly predicts the common class can look decent by accuracy while still being bad at rare but important events.

Appendix B5: Precision, Recall, and F1

F1 is built from precision and recall.

For one class, precision asks:

Of the cases the model predicted as this class, how many were actually this class?

Recall asks:

Of the cases that truly belonged to this class, how many did the model find?

For one class:

Precision⁡=true positives⁡true positives⁡+false positives⁡\operatorname{Precision} = \frac{ \operatorname{true\ positives} }{ \operatorname{true\ positives} + \operatorname{false\ positives} }

and:

Recall⁡=true positives⁡true positives⁡+false negatives⁡\operatorname{Recall} = \frac{ \operatorname{true\ positives} }{ \operatorname{true\ positives} + \operatorname{false\ negatives} }

Then:

F1=2⋅Precision⁡⋅Recall⁡Precision⁡+Recall⁡F_1 = 2 \cdot \frac{ \operatorname{Precision} \cdot \operatorname{Recall} }{ \operatorname{Precision} + \operatorname{Recall} }

A small example for one class:

true positives   = 8
false positives  = 2
false negatives  = 4

Then:

Precision⁡=88+2=0.80\operatorname{Precision} = \frac{8}{8 + 2} = 0.80

and:

Recall⁡=88+4≈0.667\operatorname{Recall} = \frac{8}{8 + 4} \approx 0.667

so:

F1=2⋅0.80⋅0.6670.80+0.667≈0.727F_1 = 2 \cdot \frac{0.80 \cdot 0.667}{0.80 + 0.667} \approx 0.727

In scikit-learn:

from sklearn.metrics import f1_score

macro_f1 = f1_score(
    y_test,
    prediction,
    average="macro"
)

weighted_f1 = f1_score(
    y_test,
    prediction,
    average="weighted"
)

Appendix B6: Macro F1 and Weighted F1

Macro F1 calculates F1 separately for each class and gives every class equal weight.

Weighted F1 also calculates F1 separately, but weights classes according to how often they occur.

Suppose a next-event model has these class-level F1 scores:

next-event class support class F1
A_Submitted 80 0.90
O_Create Offer 15 0.50
A_Cancelled 5 0.10

Macro F1 is:

Macro F1⁡=0.90+0.50+0.103=0.50\operatorname{Macro\ F1} = \frac{0.90 + 0.50 + 0.10}{3} = 0.50

Weighted F1 is:

Weighted F1⁡=80(0.90)+15(0.50)+5(0.10)100=0.80\operatorname{Weighted\ F1} = \frac{ 80(0.90) + 15(0.50) + 5(0.10) }{ 100 } = 0.80

The same model can therefore have high weighted F1 and much lower macro F1.

That pattern means:

the model is doing much better on common classes than rare classes.

The increase in macro F1:

0.604→0.6790.604 \rightarrow 0.679

therefore matters because it suggests that resource context improved prediction across the set of next-event classes, not only for the most frequent transition.

Appendix C1: Data and Reproducibility

The notebooks are published in the companion GitHub repository:

https://github.com/vbnovikov/deal-appetite-experiment

Raw datasets are not included in the repository. They should be downloaded from their original sources and placed under the local data/ directory described in DATA_SOURCES.md.

The repository is licensed under the MIT License. The datasets remain subject to their original source licenses.