# Deal Appetite Experiment

From a latent sales-ranking conjecture to stateful process prediction.

Companion repository: [vbnovikov/deal-appetite-experiment](https://github.com/vbnovikov/deal-appetite-experiment)

After looking at CRM data and business workflows across several clients, I started to wonder whether there was some underlying combination of factors that could reliably predict deal closing probability better than a simple lead score.

Things like number of touchpoints, response patterns, friction, time between interactions, stage movement, and other parts of the customer journey all seemed potentially useful when considered together.

That led to the idea that there may be a generalized ranking function over opportunities, with some part of the outcome determined by unobservable variation. If that were true, the useful question would be whether the recoverable part of that ranking could be estimated from observed CRM data, even when the full underlying process could not be observed directly.

## Short Version

The original conjecture did not survive in its broadest form.

Synthetic data showed that a recoverable appetite ranking can be learned when the structure is explicitly built into the data-generating process. Real CRM opportunity fields, early interaction aggregates, and surface-level interaction text did not provide strong or stable out-of-sample ranking of terminal outcomes.

The more useful result appeared after changing the target from distant outcome prediction to process-local prediction.

On a real event log, current state alone predicted the next event with about $0.666$ accuracy. Adding process history increased accuracy to about $0.781$. Adding resource identity increased it to about $0.838$. Adding simple graph-derived relational context increased it again to about $0.846$.

So the final claim is narrower:

> Business state may be better represented for decision-making as a history of connected entities and events than as the current values of those entities alone.

## Hypothesis

### $H_1$

There exists a recoverable ordering function over observable CRM variables:

$$
A_i = f(x_i) + \varepsilon_i
$$

such that $f(x_i)$ preserves useful relative ordering even when $\varepsilon_i$ is unobserved.

Here, $x_i$ means the feature vector for opportunity $i$: the list of measurements the model is allowed to see. Plain-language explanations of notation, feature vectors, feature groups, and preprocessing are included in [Appendix A1](#appendix-a1-libraries-and-notation), [Appendix A2](#appendix-a2-feature-vectors), [Appendix A3](#appendix-a3-feature-groups), and [Appendix A6](#appendix-a6-categorical-and-numerical-preprocessing).

### $H_0$

Observable CRM variables do not recover a stable relative ordering of opportunities beyond noise and unobserved variation.

## Experiment Setup

I generated 5,000 synthetic opportunities with four observable variables:

- fit
- urgency
- trust
- friction

Each variable was sampled independently from a standard normal distribution:

$$
f_i, u_i, t_i, r_i \sim \mathcal{N}(0,1)
$$

The observable component of appetite was defined as:

$$
S_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i
$$

I then added an unobserved term:

$$
\varepsilon_i \sim \mathcal{N}(0,1.5^2)
$$

giving total latent appetite:

$$
A_i = S_i + \varepsilon_i
$$

$A_i$ was converted to a closing probability using the logistic function:

$$
p_i = \frac{1}{1 + e^{-A_i}}
$$

and the observed outcome was sampled as:

$$
Y_i \sim \operatorname{Bernoulli}(p_i)
$$

A logistic regression was then trained using only $f_i$, $u_i$, $t_i$, and $r_i$. The latent appetite, unobserved term, and true closing probability were withheld from the model.

The point of the synthetic setup is to separate what is observable from what is hidden. [Appendix A2](#appendix-a2-feature-vectors) explains how observable columns become model inputs, and [Appendix B2](#appendix-b2-roc-auc) and [Appendix B3](#appendix-b3-spearman-rank-correlation) explain the ranking metrics used to evaluate the model.

## Evaluation

The model was evaluated against three different targets.

The metrics used here are defined in plain language in [Appendix B2](#appendix-b2-roc-auc) and [Appendix B3](#appendix-b3-spearman-rank-correlation).

### 1. Observed outcome

How well does the model rank deals that actually closed above deals that did not?

This is measured using ROC AUC against $Y_i$.

### 2. Total latent appetite

How closely does the model's ranking agree with the full latent quantity:

$$
A_i = S_i + \varepsilon_i
$$

This includes the unobserved term, so perfect recovery should not be possible.

### 3. Systematic appetite

How closely does the model recover the ordering generated by the observable component:

$$
S_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i
$$

This is the part of the synthetic process that should, in principle, be recoverable from the available data.

## Results

The model achieved an ROC AUC of approximately:

$$
\operatorname{AUC} \approx 0.811
$$

against the observed close/no-close outcome.

That was encouraging, but the more important comparison was against the quantities that are known only because the data is synthetic.

The model had limited agreement with total latent appetite,

$$
A_i = S_i + \varepsilon_i
$$

because $\varepsilon_i$ is deliberately hidden:

$$
\rho_s(\hat{p}, A) \approx 0.791
$$

Its ranking of the systematic component $S_i$, however, was nearly perfect:

$$
\rho_s(\hat{p}, S) \approx 0.999
$$

In the synthetic setting, the recoverable part of the ranking function was therefore recoverable from the observable variables.

## The Problem

The synthetic experiment does not establish that a comparable latent variable exists in real CRM data.

The systematic component was defined directly from the same variables given to the model:

$$
S_i = 1.2f_i + 0.9u_i + 0.7t_i - 1.0r_i
$$

A model recovering that ordering therefore shows that the estimation procedure works under the assumptions of the simulation. It does not show that a comparable one-dimensional quantity exists in real CRM data.

The unobserved term creates another problem. In a real business process, $\varepsilon_i$ is not a clean random variable whose distribution is known. It can contain missing interactions, individual behavior, timing, relationships, information outside the CRM, and variables that were never recorded.

At that point, the conjecture loses a clear interpretation because the unobserved component can contain almost anything the model does not capture.

## Real CRM Data

The next experiment used a real CRM sales dataset with resolved opportunities linked to account, product, and sales-team data.

The model only used information that would have been known while the opportunity was still open. Fields like the final close date or realized deal value were excluded because they would leak the outcome into the model and make the test meaningless.

To avoid leaking future information into the training set, the data was split chronologically:

$$
\text{Train} = \{i : t_i < T\}
$$

$$
\text{Test} = \{i : t_i \geq T\}
$$

The model was trained on earlier opportunities and evaluated on later ones.

This is a temporal holdout: the test cases occur later in time than the training cases. [Appendix A7](#appendix-a7-training-only-preprocessing) explains why the same boundary also matters for preprocessing.

The baseline result was:

$$
\operatorname{AUC} \approx 0.526
$$

which is only slightly above random ranking.

### Feature Ablation

I then removed and added groups of variables while keeping the model and temporal split fixed.

The results were:

$$
\operatorname{AUC}_{\text{Core}} \approx 0.530
$$

$$
\operatorname{AUC}_{\text{Core + Rep}} \approx 0.540
$$

$$
\operatorname{AUC}_{\text{Core + Rep + Account}} \approx 0.525
$$

Salesperson and team information added a small amount of predictive value. Adding account identity reduced performance on the future holdout set.

The static CRM fields available in this dataset therefore provided very little stable out-of-sample ranking power.

## Interaction History

The weak performance of the static CRM model raised a more specific question: does the history of interactions with an opportunity contain useful information that is missing from the opportunity record itself?

The next experiment began with a comparison between static CRM data and static CRM data plus interaction-derived information. In practice, the available interaction records introduced a narrower test: whether coarse first-30-day activity summaries and surface-level interaction text contained stable ranking signal under temporal validation.

### Static CRM Model

The first model estimates the probability that opportunity $i$ will be won using only static CRM features:

$$
\hat{p}_i^{(s)}
=
P(Y_i = 1 \mid \mathbf{x}_i^{(s)})
$$

where $\mathbf{x}_i^{(s)}$ is the feature vector containing the static CRM information supplied to the model.

If this notation is unfamiliar, read $\mathbf{x}_i^{(s)}$ as "the static-data row for opportunity $i$ after it has been converted into numbers." [Appendix A1](#appendix-a1-libraries-and-notation) explains the notation, and [Appendix A2](#appendix-a2-feature-vectors) gives concrete examples.

The static feature set includes fields such as:

- product
- account
- sector
- salesperson
- regional office
- company revenue
- number of employees
- deal value
- engagement month

Categorical fields are one-hot encoded and numerical fields are standardized using statistics from the training set. The preprocessing details are included in [Appendix A6](#appendix-a6-categorical-and-numerical-preprocessing) and [Appendix A7](#appendix-a7-training-only-preprocessing).

### Interaction-Derived Features

A natural next model would include information derived from the communication history associated with the opportunity:

$$
\hat{p}_i^{(s+b)}
=
P\left(
Y_i = 1
\mid
\mathbf{x}_i^{(s)},
\mathbf{x}_i^{(b)}
\right)
$$

The new term,

$$
\mathbf{x}_i^{(b)}
$$

is another feature vector, this time derived from interactions rather than static CRM fields.

The superscripts here label feature groups. They are not exponents. $\mathbf{x}_i^{(s)}$ means static features, and $\mathbf{x}_i^{(b)}$ means behavioral or interaction-derived features.

Depending on the interaction record, these features can describe things such as:

- number of interactions
- time between interactions
- recency of communication
- direction of communication
- who initiated the interaction
- interaction type
- properties extracted from communication text

For example, an opportunity might have the following interaction history:

```text
Day 1   Salesperson → Customer
Day 3   Customer → Salesperson
Day 8   Salesperson → Customer
Day 9   Customer → Salesperson
```

That sequence can be converted into numerical features such as:

```text
interaction_count       = 4
customer_replies        = 2
salesperson_messages    = 2
days_since_last_contact = 1
mean_response_delay     = ...
```

Those values can then be encoded into $\mathbf{x}_i^{(b)}$ and supplied to the model alongside the static feature vector.

The combined expression can therefore be read as:

> the probability that opportunity $i$ is won, given both its static CRM data and its interaction history.

The motivating comparison is whether the combined representation improves on static CRM data alone:

$$
\operatorname{AUC}_{s+b}
>
\operatorname{AUC}_{s}
$$

However, the available interaction records were not clean opportunity-level histories. Many interactions were better interpreted as relationship-level observations between salesperson and customer contact. The implemented baselines therefore tested a narrower question first: whether first-30-day activity summaries or first-30-day lexical content contained stable ranking signal under temporal validation.

Implementation details for preprocessing, modeling, and evaluation are included in the appendices, especially [Appendix A6](#appendix-a6-categorical-and-numerical-preprocessing), [Appendix A7](#appendix-a7-training-only-preprocessing), and [Appendix B1](#appendix-b1-evaluation-tasks).

## Interaction History Results

The implemented interaction-history baselines tested:

$$
\operatorname{AUC}_{\text{activity}}
$$

for first-30-day activity summaries, and:

$$
\operatorname{AUC}_{\text{text}}
$$

for first-30-day TF-IDF text features.

The initial interaction baselines did not find useful out-of-sample ranking signal. The activity-only temporal holdout produced:

$$
\operatorname{AUC}_{\text{activity}} \approx 0.473
$$

and the text-only baseline produced:

$$
\operatorname{AUC}_{\text{text}} \approx 0.495
$$

Both are at or below the random-ranking benchmark:

$$
\operatorname{AUC} = 0.5
$$

This does not imply that interaction history is irrelevant. It shows that these initial representations did not recover stable outcome-ranking information in this dataset.

One likely issue is compression. Interaction history contains more detail than the static opportunity record, but much of that detail is lost before it reaches the model.

For example, a sequence such as:

```text
Day 1   Salesperson → Customer
Day 2   Customer → Salesperson
Day 3   Customer → Salesperson
Day 12  Salesperson → Customer
```

might be reduced to features such as:

```text
interaction_count       = 4
customer_messages       = 2
salesperson_messages    = 2
mean_response_delay     = ...
days_since_last_contact = ...
```

Those aggregate values preserve some information about the interaction history, but they no longer describe the sequence itself.

Two opportunities can therefore have similar aggregate features while having very different histories.

For example:

```text
Opportunity A

Salesperson → Customer
Customer → Salesperson
Salesperson → Customer
Customer → Salesperson
```

and:

```text
Opportunity B

Salesperson → Customer
Salesperson → Customer
Customer → Salesperson
Customer → Salesperson
```

can produce similar counts even though the order of events is different.

The same issue applies to timing. Four interactions spread evenly across a week describe a different process from four interactions followed by a month of silence, even if both opportunities have the same total interaction count.

This led to the next question: Is the useful information contained in the trajectory of an opportunity rather than in aggregate features calculated from that trajectory?

## Sequential Customer Journeys

The interaction-history experiment reduced a sequence of events to summary variables such as interaction count, recency, and response timing. That loses information about the order in which those events occurred.

The next experiment tested whether preserving more of the sequence improves prediction.

Instead of treating an opportunity as one static feature vector, each customer journey is represented as an ordered series of events.

This is the sequence version of the same idea. A feature vector stores a fixed list of inputs for one example; a sequence stores an ordered list of events for one example. [Appendix A2](#appendix-a2-feature-vectors) gives the feature-vector version first, and [Appendix A5](#appendix-a5-histories-as-feature-vectors) shows how histories are converted into model inputs.

For session $i$, let:

$$
E_i
=
\left(
e_{i1},
e_{i2},
\ldots,
e_{in_i}
\right)
$$

where:

- $E_i$ is the complete observed journey for session $i$
- $e_{i1}$ is the first event
- $e_{i2}$ is the second event
- $n_i$ is the total number of events observed in that session

A simple journey might look like:

```text
home
→ product_page
→ cart
→ checkout
→ confirmation
```

The order is important because a customer who reaches `checkout` has followed a different process from someone whose session ends on the product page, even if both generated the same number of events.

### Journey Prefixes

The model should only use information that would have been available at the time a prediction was made.

For that reason, the experiment does not give the model the complete journey at once. It constructs progressively longer prefixes of the journey.

The first $k$ events are represented as:

$$
E_i^{(k)}
=
\left(
e_{i1},
e_{i2},
\ldots,
e_{ik}
\right)
$$

For example:

```text
k = 1

home
```

```text
k = 2

home
→ product_page
```

```text
k = 3

home
→ product_page
→ cart
```

```text
k = 4

home
→ product_page
→ cart
→ checkout
```

At each point, the model estimates:

$$
\hat{p}_i^{(k)}
=
P
\left(
Y_i = 1
\mid
E_i^{(k)}
\right)
$$

Here:

- $Y_i = 1$ means that session $i$ eventually results in a purchase
- $E_i^{(k)}$ is everything observed in that session up to event $k$
- $\hat{p}_i^{(k)}$ is the estimated probability of eventual conversion based only on that partial journey

The expression can be read as:

> the estimated probability that session $i$ eventually converts, given the first $k$ events observed in its customer journey.

The experiment then compares:

$$
\operatorname{AUC}(1),
\operatorname{AUC}(2),
\operatorname{AUC}(3),
\operatorname{AUC}(4)
$$

If the ordering becomes more accurate as $k$ increases, additional process history is providing useful information about the eventual outcome.

### Avoiding Outcome Leakage

The dataset contains a `Purchased` field on every row in a session.

That creates an important problem:

If a session eventually results in a purchase, earlier rows in the same session can also contain:

```text
Purchased = 1
```

even though the purchase had not happened yet.

Using that field as an input would effectively tell the model the answer in advance.

The `confirmation` event creates a similar issue. Reaching confirmation means the purchase has already occurred, so including it in the predictive history would make the experiment trivial.

For that reason:

- `Purchased` is used only as the final target
- `confirmation` is treated as a post-outcome event
- only events that occur before confirmation are used as model inputs

This keeps the question meaningful: given only what had happened so far, could the model rank sessions by their eventual probability of conversion?

This is the same leakage rule used throughout the experiments: do not let information from after the prediction time enter the inputs. [Appendix A7](#appendix-a7-training-only-preprocessing) describes the rule in preprocessing terms.

### Constructing the Sequences

The event data is first ordered by session and timestamp:

```python
customer_journey["Timestamp"] = pd.to_datetime(
    customer_journey["Timestamp"]
)

customer_journey = (
    customer_journey
    .sort_values(["SessionID", "Timestamp"])
    .reset_index(drop=True)
)
```

Each event is then numbered within its session:

```python
customer_journey["event_number"] = (
    customer_journey
    .groupby("SessionID")
    .cumcount()
    + 1
)
```

A session such as:

```text
Session 42

12:01   home
12:03   product_page
12:07   cart
12:10   checkout
```

therefore becomes:

```text
event_number   page
1              home
2              product_page
3              cart
4              checkout
```

This makes it possible to construct the first event, first two events, first three events, and so on without using anything that occurred later in the session.

The important change in this experiment is that the prediction is now conditioned on an ordered process history rather than only on a static record or a set of aggregate interaction features.

## Sequential Journey Results

Ranking performance did not improve meaningfully as more of the customer journey became visible.

The pre-confirmation prefix models produced:

$$
\operatorname{AUC}(1) \approx 0.457
$$

$$
\operatorname{AUC}(2) \approx 0.464
$$

$$
\operatorname{AUC}(3) \approx 0.459
$$

$$
\operatorname{AUC}(4) \approx 0.450
$$

All four results are below the random-ranking benchmark:

$$
\operatorname{AUC} = 0.5
$$

So this experiment did not support the idea that these customer-journey prefixes recover a stable ranking of eventual conversion.

It did, however, expose an important interpretation problem for any future sequence experiment.

Consider these two partial journeys:

```text
Session A

home
→ product_page
→ cart
```

```text
Session B

home
→ product_page
→ cart
```

Both sessions have the same current state: `cart`.

If a model distinguishes sessions mainly because one has reached `cart` while another remains on `home`, then it has learned funnel position. That does not show that the path taken to reach the current state contains additional information.

The distinction can be written as:

$$
P(Y_i = 1 \mid S_i)
$$

versus:

$$
P(Y_i = 1 \mid S_i, H_i)
$$

where:

- $S_i$ is the current state of session $i$
- $H_i$ is the history that led to that state
- $Y_i = 1$ means the session eventually converts

The first expression asks:

> What is the probability of conversion given where the session is now?

The second asks:

> What is the probability of conversion given both where the session is now and how it got there?

That is a much stronger test.

### Current State as a Baseline

Suppose one customer has reached `checkout` after moving directly through the funnel:

```text
home
→ product_page
→ cart
→ checkout
```

Another customer is also currently at `checkout`, but arrived there after repeated navigation between earlier states:

```text
home
→ product_page
→ home
→ product_page
→ cart
→ product_page
→ cart
→ checkout
```

A conventional CRM-style snapshot could represent both sessions as:

```text
current_state = checkout
```

The current-state model therefore receives the same main piece of information for both.

A history-aware representation can preserve the different trajectories.

The relevant comparison becomes:

$$
\operatorname{Performance}(S_i, H_i)
>
\operatorname{Performance}(S_i)
$$

If adding $H_i$ improves prediction while $S_i$ is held constant, then the trajectory itself contains information that is lost in the current-state representation.

This became the basis for the next experiments.

Put simply, the current state is like a summary column, while the history is the set of steps that produced that summary. The next experiments ask whether that summary column has thrown away useful information.

## State Transitions

To isolate the effect of history, the next experiment uses an explicit state-transition process.

Instead of predicting only whether a deal eventually closes, the process is represented as a sequence of states:

$$
S_{i,0}
\rightarrow
S_{i,1}
\rightarrow
S_{i,2}
\rightarrow
\cdots
\rightarrow
S_{i,t}
$$

Here, $S_{i,t}$ is the state of process $i$ at time $t$.

A simplified sales process might look like:

```text
New
→ Contacted
→ Qualified
→ Proposal
→ Won
```

but another opportunity might follow:

```text
New
→ Contacted
→ Qualified
→ Contacted
→ Qualified
→ Proposal
→ Lost
```

At some point, both opportunities can occupy the same state even though their histories are different.

That gives us a cleaner question:

> Among opportunities that are currently in the same state, does their previous transition history help predict what happens next?

Formally, the baseline uses only the current state:

$$
P(S_{i,t+1} \mid S_{i,t})
$$

The history-aware model uses the current state and the preceding trajectory:

$$
P(
S_{i,t+1}
\mid
S_{i,t},
S_{i,t-1},
\ldots,
S_{i,0}
)
$$

The first expression asks for the probability of the next state given only the current state.

The second asks whether knowing the path taken to reach that state changes the probability of what happens next.

This removes much of the ambiguity from the customer-journey experiment. The comparison is no longer between customers at obviously different points in a funnel. It compares cases that occupy the same current state and asks whether their histories still matter.

That is why the later experiments are stricter than ordinary funnel-stage prediction. They compare like with like: same current state, different prior path.

## Within-State Transition History

The previous experiment showed that later process states are easier to rank, but that does not tell us whether the history leading to a state matters.

To test that directly, I built a synthetic state-transition process where opportunities move through a small sales funnel:

$$
\mathcal{S}
=
\{
\text{new},
\text{engaged},
\text{qualified},
\text{proposal},
\text{won},
\text{lost}
\}
$$

Each opportunity starts in `new` and moves through the process one step at a time.

At any non-terminal state, it can:

- advance
- remain in the same state
- regress
- become lost

For example:

```text
new
→ engaged
→ qualified
→ proposal
→ won
```

is one possible trajectory, while:

```text
new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal
→ lost
```

is another.

### A Hidden Propensity

Each synthetic opportunity is assigned an unobserved value:

$$
A_i
$$

which affects how likely it is to move forward, stall, regress, or become lost.

Higher values of $A_i$ make forward movement more likely and loss less likely.

The transition process can be written as:

$$
P
\left(
S_{i,t+1}
\mid
S_{i,t},
A_i
\right)
$$

where:

- $S_{i,t}$ is the current state of opportunity $i$
- $S_{i,t+1}$ is its next state
- $A_i$ is the hidden propensity assigned to that opportunity

The expression can be read as:

> the probability of the opportunity's next state, given its current state and its unobserved propensity.

The model never receives $A_i$. It only observes what happened to the opportunity.

### Generating Transition Probabilities

Each possible transition is first assigned a score.

For example:

$$
z_{\text{advance}}
=
\alpha_s + \beta_a A_i
$$

$$
z_{\text{stay}}
=
0
$$

$$
z_{\text{regress}}
=
\gamma_s - \beta_r A_i
$$

$$
z_{\text{lost}}
=
\lambda_s - \beta_l A_i
$$

The coefficients control how strongly the hidden propensity affects each type of transition.

These scores are not probabilities yet. They can take any real value.

To convert them into probabilities that add up to $1$, the experiment uses the softmax function:

$$
P(T_i = j)
=
\frac{
\exp(z_{ij})
}{
\sum_{m \in \mathcal{T}(S_{i,t})}
\exp(z_{im})
}
$$

Here, $\mathcal{T}(S_{i,t})$ is the set of transitions that are available from the current state.

For example, an opportunity in `qualified` might be able to:

```text
advance → proposal
stay    → qualified
regress → engaged
exit    → lost
```

Softmax converts the four transition scores into four probabilities whose total is:

$$
1
$$

One of those transitions is then sampled, producing the next observed state.

This process repeats until the opportunity reaches either `won` or `lost`.

### Why Use Synthetic Data Again?

The question is no longer whether a general latent deal score exists.

The synthetic process gives us a controlled environment where we know that past transitions contain information about an unobserved variable.

That lets us test a specific representation question:

> If two opportunities are currently in the same state, can their previous transitions help distinguish them?

Suppose two opportunities are both currently `qualified`.

The first followed:

```text
new
→ engaged
→ qualified
```

The second followed:

```text
new
→ engaged
→ qualified
→ engaged
→ qualified
```

A current-state representation gives both of them:

```text
current_state = qualified
```

A history-aware representation can distinguish the two trajectories.

The comparison is therefore:

$$
P(Y_i = 1 \mid S_{i,t})
$$

against:

$$
P(Y_i = 1 \mid S_{i,t}, H_{i,t})
$$

where $H_{i,t}$ contains the transitions observed before time $t$.

If the history-aware model performs better among opportunities occupying the same current state, then current state is losing information about the process that produced it.

That is the specific property being tested here.

### Within-State Ranking Results

The history-aware model was evaluated separately within the `engaged`, `qualified`, and `proposal` states.

The results were:

$$
\operatorname{AUC}_{\text{engaged}} \approx 0.647
$$

$$
\operatorname{AUC}_{\text{qualified}} \approx 0.630
$$

$$
\operatorname{AUC}_{\text{proposal}} \approx 0.666
$$

All three are above the random-ranking benchmark:

$$
\operatorname{AUC} = 0.5
$$

Because each model is evaluated only among opportunities that occupy the same current state, differences in pipeline stage cannot explain the ranking.

For example, every opportunity in the `proposal` evaluation set is already at `proposal`. The model therefore cannot perform well simply by learning that `proposal` is generally closer to `won` than `engaged`.

The information available to the model comes from the history that produced that state, including features such as:

- number of forward transitions
- number of regressions
- number of repeated states
- total number of transitions
- net progress
- regression rate
- most recent transition

So, within the synthetic process, two opportunities can both be at:

```text
proposal
```

while their histories imply different probabilities of eventually reaching `won`.

For example:

```text
Opportunity A

new
→ engaged
→ qualified
→ proposal
```

and:

```text
Opportunity B

new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal
```

have the same current state but different trajectories.

The result shows that the trajectory contains information that is lost when both opportunities are represented only as:

```text
current_state = proposal
```

### Limitation

This result still comes from a synthetic process.

The hidden propensity $A_i$ was deliberately constructed to affect transition probabilities. An opportunity with a higher $A_i$ is more likely to advance and less likely to regress or become lost.

Transition history is therefore expected to contain information about $A_i$.

The result does not establish that real sales or business processes behave this way. It shows that if an unobserved factor affects how entities move through a process, their observed transition history can preserve information about that factor even after their current state is known.

That gives us a real-data question:

> among cases observed at the same point in a real business process, does the history leading to that point contain useful information about what eventually happens?
## Testing Transition History on Real Process Data

The synthetic state-transition experiment shows that history can matter when the process is explicitly constructed so that previous transitions contain information about an unobserved variable.

That still does not tell us whether the same thing happens in real business processes.

The next experiment uses the BPI Challenge 2017 event log, which contains event-level histories for real loan applications. Unlike a normal CRM snapshot, the dataset records the sequence of activities associated with each application over time.

A simplified application history might look like:

```text
A_Create Application
→ A_Submitted
→ A_Concept
→ A_Accepted
→ O_Create Offer
→ O_Sent
→ A_Pending
```

Another application may reach some of the same intermediate states through a different sequence.

That makes it possible to repeat the within-state test using real process history.

### Defining the Outcome

Each application eventually reaches one of three terminal states:

```text
A_Pending
A_Denied
A_Cancelled
```

For the binary experiment, successful completion is defined as:

$$
Y_i
=
\begin{cases}
1 & \text{if application } i \text{ ends in A\_Pending} \\
0 & \text{if application } i \text{ ends in A\_Denied or A\_Cancelled}
\end{cases}
$$

The target is determined only from the final terminal event.

Intermediate activities such as `A_Accepted` are not treated as successful outcomes because the application can still fail later in the process.

This distinction between intermediate events and final labels is another form of leakage control. The model is not allowed to treat a promising middle state as if it were already the final answer.

### Comparing Applications at the Same State

The key requirement is that the applications being compared must occupy the same observable process state.

Suppose two applications have both reached:

```text
O_Sent
```

At that point, their current state is identical.

For a checkpoint state $s$, define:

$$
\mathcal{C}_s
=
\{
i :
\text{application } i \text{ reaches state } s
\}
$$

$\mathcal{C}_s$ is simply the set of applications that reached the same checkpoint.

For every application in that set, the model is allowed to use only the events that occurred before or at the first occurrence of that checkpoint.

Anything that happens later is hidden.

This gives each application a history:

$$
H_{i,t}
=
\left(
S_{i,0},
S_{i,1},
\ldots,
S_{i,t}
\right)
$$

while holding the final observed state fixed:

$$
S_{i,t} = s
$$

The question is then:

$$
P(Y_i = 1 \mid S_{i,t}=s, H_{i,t})
$$

Can the history $H_{i,t}$ help rank applications that are all currently at the same state $s$?

If it can, then the current state alone is not a sufficient description of the process.

### Why the First Occurrence Matters

Some applications can return to the same activity more than once.

For example:

```text
A_Submitted
→ A_Concept
→ A_Accepted
→ A_Concept
→ A_Accepted
```

If we simply chose an arbitrary occurrence of `A_Accepted`, some applications would have much more future information available than others.

The experiment therefore uses the first occurrence of the chosen checkpoint.

Everything after that point is excluded.

This gives every application the same prediction condition:

> what could have been known the first time this application reached this state?

### Temporal Validation

The data is again split chronologically.

Applications that started earlier are used for training, while later applications are reserved for testing.

This matters because a random split could allow the model to learn process patterns from future cases and use them to predict earlier ones.

The real question is whether historical process data can help rank applications that occur later in time.

### Within-State AUC

For each checkpoint state $s$, ranking performance is measured separately:

$$
\operatorname{AUC}_s
$$

Because every application in that evaluation set already occupies the same current state, the state itself cannot explain differences in ranking.

A value of:

$$
\operatorname{AUC}_s = 0.5
$$

would mean the history provides no useful ordering within that state.

A value above:

$$
0.5
$$

would indicate that the path leading to the checkpoint contains information about the eventual terminal outcome.

That is a much stronger test than asking whether applications further along in the process are more likely to succeed.

### Real Within-State Results

The first history-only models used aggregate process features such as event counts, elapsed time, repeated activities, offer activity, and workflow activity.

The `A_Submitted` checkpoint was excluded before modeling because almost every application reached it through the same observed history. There was essentially no variation to rank on.

The remaining checkpoints were:

- `A_Validating`
- `A_Incomplete`

The history-only models produced:

$$
\operatorname{AUC}_{A\_Validating} \approx 0.581
$$

and:

$$
\operatorname{AUC}_{A\_Incomplete} \approx 0.589
$$

Both are above the random-ranking benchmark:

$$
\operatorname{AUC} = 0.5
$$

Because each model is evaluated only among applications at the same checkpoint, the result cannot be explained by one application simply being further through the process than another.

The difference comes from what happened before the checkpoint.

The effect is still fairly weak. An AUC around $0.58$ does not support accurate within-state ranking, but it does suggest that some information about the eventual outcome survives in the prior process history.

### Preserving More of the Sequence

The first real-data model still compresses history into aggregate counts.

For example, these histories:

```text
A → B → C → B → C
```

and:

```text
A → B → B → C → C
```

can produce similar event counts despite having different transition structures.

The next version therefore added features describing adjacent transitions:

$$
(e_{i1} \rightarrow e_{i2}),
(e_{i2} \rightarrow e_{i3}),
\ldots,
(e_{i,n-1} \rightarrow e_{in})
$$

Instead of recording only how often an activity occurred, the model could also record transitions such as:

```text
A_Validating → A_Incomplete
A_Incomplete → W_Call incomplete files
O_Create Offer → O_Sent
```

The sequence-aware model produced:

$$
\operatorname{AUC}_{A\_Validating} \approx 0.592
$$

and:

$$
\operatorname{AUC}_{A\_Incomplete} \approx 0.568
$$

Compared with the aggregate-history model:

$$
0.581 \rightarrow 0.592
$$

for `A_Validating`, while:

$$
0.589 \rightarrow 0.568
$$

for `A_Incomplete`.

Preserving adjacent transition structure therefore improved one checkpoint slightly and made the other worse.

There is some recoverable information in the process history, but simply adding more sequence detail does not produce a consistently stronger ranking model.

### What This Actually Supports

The original conjecture was much broader: that opportunity outcomes might be recoverable through a generalized ranking function representing some latent transaction propensity.

These experiments do not provide strong evidence for that claim.

The real-data results are much narrower. Applications observed at the same process state can still differ in ways that are partially recoverable from the history that brought them there.

A useful way to state that is:

$$
S_{i,t}
=
g(H_{i,t})
$$

where the current state $S_{i,t}$ is produced by some preceding history $H_{i,t}$.

Representing only $S_{i,t}$ therefore compresses that history into a much smaller description of the process.

Different histories can map to the same current state:

$$
H_a \neq H_b
$$

while:

$$
g(H_a) = g(H_b)
$$

Some information that distinguishes $H_a$ from $H_b$ is necessarily absent from the current-state representation.

The empirical question is whether any of that discarded information matters for the decision being made.

In the BPI experiment, at least some of it does.

The effect is modest, but the result is enough to motivate a different prediction target: rather than trying to predict a distant terminal outcome such as `won`, `lost`, or successful completion, can the current process history predict what happens next?

## Next-State Prediction

The terminal-outcome experiments produced only modest within-state ranking performance. The next experiment changes the target.

Instead of asking whether an application will eventually succeed, the model predicts the **next observable process event**.

For case $i$ at time $t$, let:

$$
S_{i,t}
$$

represent the current state,

$$
H_{i,t}
$$

the process history observed so far, and:

$$
G_{i,t}
$$

the relational context surrounding the case.

The target is:

$$
S_{i,t+1}
$$

so the full prediction problem becomes:

$$
P(
S_{i,t+1}
\mid
S_{i,t},
H_{i,t},
G_{i,t}
)
$$

This can be read as:

> the probability of the next process state, given the current state, the history so far, and the surrounding relational context.

The notation is compact, but the idea is ordinary supervised learning: each row contains what was known at time $t$, and the label is what happened at time $t+1$. [Appendix A5](#appendix-a5-histories-as-feature-vectors) translates this into a simple table-style example.

### Constructing Prediction Snapshots

The BPI event log is first ordered by application and timestamp:

```python
bpi = (
    bpi
    .sort_values(
        ["case:concept:name", "time:timestamp"]
    )
    .reset_index(drop=True)
)
```

Each event is then assigned its position within the application:

```python
bpi["event_position"] = (
    bpi
    .groupby("case:concept:name")
    .cumcount()
)
```

The next event is created by shifting the activity column one row backward within each application:

```python
bpi["next_activity"] = (
    bpi
    .groupby("case:concept:name")["concept:name"]
    .shift(-1)
)
```

This turns an event sequence such as:

```text
A_Submitted
→ A_Concept
→ A_Accepted
→ O_Create Offer
```

into supervised examples:

```text
current: A_Submitted
target:  A_Concept
```

```text
current: A_Concept
target:  A_Accepted
```

```text
current: A_Accepted
target:  O_Create Offer
```

The last event in each application is removed because there is no following event to predict:

```python
event_snapshots = bpi[
    bpi["next_activity"].notna()
].copy()
```

### Temporal Train/Test Split

The split is performed at the application level using case start time.

First, the start time of each application is calculated:

```python
case_start_times = (
    bpi
    .groupby("case:concept:name")["time:timestamp"]
    .min()
)
```

A chronological cutoff is then chosen:

```python
cutoff = case_start_times.quantile(0.75)
```

Earlier applications form the training set:

```python
train_snapshots = event_snapshots[
    event_snapshots["case_start_time"] < cutoff
].copy()
```

Later applications form the test set:

```python
test_snapshots = event_snapshots[
    event_snapshots["case_start_time"] >= cutoff
].copy()
```

This preserves the same basic rule used throughout the experiments: the model should learn from the past and be evaluated on later cases.

[Appendix A7](#appendix-a7-training-only-preprocessing) gives the general version of this rule, and [Appendix B1](#appendix-b1-evaluation-tasks), [Appendix B4](#appendix-b4-accuracy), and [Appendix B6](#appendix-b6-macro-f1-and-weighted-f1) explain why the reported next-event metrics are accuracy, macro F1, and weighted F1 rather than only AUC.

## State-Only Baseline

The first baseline uses only the current activity.

For every state $s$, the training data is used to find the next event that followed that state most often:

$$
\hat{S}_{t+1}(s)
=
\arg\max_j
P(S_{t+1}=j \mid S_t=s)
$$

The $\arg\max$ operator means:

> choose the value of $j$ that gives the largest probability.

In code:

```python
state_next_mode = (
    train_snapshots
    .groupby("concept:name")["next_activity"]
    .agg(
        lambda x: x.value_counts().idxmax()
    )
)
```

If `A_Concept` was most often followed by `A_Accepted` in the historical training data, then the baseline predicts:

```text
A_Concept → A_Accepted
```

whenever it encounters `A_Concept`.

Those predictions are applied to the test set:

```python
test_snapshots["state_only_prediction"] = (
    test_snapshots["concept:name"]
    .map(state_next_mode)
)
```

The resulting temporal accuracy was:

$$
\operatorname{Accuracy}_{\text{state-only}}
\approx
0.666
$$

with complete coverage of the test set.

So current state alone correctly predicts the next event about two-thirds of the time.

That is already much stronger than the earlier attempts to predict a distant terminal outcome.

## Adding Process History

The next model includes compact information about the path observed so far.

The history representation includes:

- event position
- elapsed time since application creation
- previous activity
- number of unique activities observed
- number of repeated activities
- cumulative application-event count
- cumulative offer-event count
- cumulative workflow-event count

The prediction changes from:

$$
P(S_{t+1} \mid S_t)
$$

to:

$$
P(S_{t+1} \mid S_t, H_t)
$$

A simplified feature set might look like:

```text
current_activity        = A_Validating
previous_activity       = A_Incomplete
event_position          = 12
elapsed_hours           = 47.3
unique_activity_count   = 8
repeated_activity_count = 4
offer_event_count       = 2
workflow_event_count    = 5
```

Categorical features such as `current_activity` and `previous_activity` are one-hot encoded. Numerical features are standardized. [Appendix A6](#appendix-a6-categorical-and-numerical-preprocessing) shows why both transformations are needed before a standard classifier can read the data.

The model is built with scikit-learn:

```python
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.linear_model import SGDClassifier
```

The preprocessing step can be defined as:

```python
preprocessor = ColumnTransformer(
    transformers=[
        (
            "categorical",
            OneHotEncoder(handle_unknown="ignore"),
            categorical_features,
        ),
        (
            "numeric",
            StandardScaler(),
            numeric_features,
        ),
    ]
)
```

The classifier and preprocessing are then combined into a single pipeline:

```python
model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        (
            "classifier",
            SGDClassifier(
                loss="log_loss",
                max_iter=2000,
                random_state=42,
            ),
        ),
    ]
)
```

The important point is that preprocessing is learned only from the training data.

The model is then fit with:

```python
model.fit(
    X_train,
    y_train
)
```

and next-event predictions are generated for the temporal holdout:

```python
history_prediction = model.predict(
    X_test
)
```

The resulting accuracy was:

$$
\operatorname{Accuracy}_{\text{state+history}}
\approx
0.781
$$

Compared with the state-only baseline:

$$
0.666
\rightarrow
0.781
$$

which is an absolute improvement of:

$$
0.115
$$

or about **11.5 percentage points**.

The history of the process is therefore much more useful for predicting the next event than it was for predicting a distant final outcome.

## Relational Context Ablation

The next step was to test which additional forms of context actually mattered.

Rather than adding every available field at once, the experiment uses **ablation**.

Ablation means training several otherwise similar models while adding or removing specific groups of features. If performance changes materially when one group is introduced, that group is contributing useful predictive information.

The main representations were:

```text
State + History
State + History + Case
State + History + Resource
State + History + Case + Resource
```

The models were evaluated with:

- accuracy
- macro F1
- weighted F1

These are classification metrics rather than ranking metrics because the target is now a specific next event, not a binary win/loss outcome. [Appendix B1](#appendix-b1-evaluation-tasks) explains the difference.

The results were:

| Representation | Accuracy | Macro F1 | Weighted F1 |
|---|---:|---:|---:|
| State + History | 0.781 | 0.604 | 0.751 |
| State + History + Case | 0.786 | 0.613 | 0.757 |
| State + History + Resource | 0.838 | 0.679 | 0.820 |
| State + History + Case + Resource | 0.840 | 0.683 | 0.823 |

Adding case-level information produced only a small improvement.

Adding resource identity produced a much larger one:

$$
0.781
\rightarrow
0.838
$$

### What Does "Resource" Mean?

In the BPI process log, events can be associated with the resource responsible for that activity.

That might represent a user, employee, team member, or operational handler.

So a row can contain information conceptually like:

```text
case      = Application_4821
activity  = A_Validating
resource  = User_37
```

When resource identity is added to the model, it can learn that different resources are associated with different transition patterns.

That does **not** establish that the resource causes the transition.

Resource identity can also capture:

- specialization
- role
- assignment rules
- case complexity
- organizational structure
- operational routing
- workload patterns

The ablation establishes only that this information helps predict what happens next. Metric definitions for accuracy, macro F1, and weighted F1 are included in [Appendix B4](#appendix-b4-accuracy), [Appendix B5](#appendix-b5-precision-recall-and-f1), and [Appendix B6](#appendix-b6-macro-f1-and-weighted-f1).

## Graph-Derived Context

Resource identity is relational information, but encoding a resource as a category still treats it as a label.

For example:

```text
resource = User_37
```

does not tell the model how `User_37` is connected to other applications or resources.

The event log can also be represented as a graph.

A simplified local structure might look like:

```text
Application A
├── handled by → Resource 12
├── handled by → Resource 41
└── contains   → Event ...

Application B
├── handled by → Resource 12
└── contains   → Event ...
```

Because `Resource 12` appears in both applications, the two cases are connected through a shared node.

Graph-derived features summarize parts of this local structure and add them to the existing state, history, and resource representation.

At this point, the feature vector is no longer just a set of fields from one row. It also includes simple summaries of the case's local graph context.

The final comparison was:

| Representation | Accuracy | Macro F1 | Weighted F1 |
|---|---:|---:|---:|
| State only | 0.666 | 0.483 | 0.609 |
| State + History | 0.781 | 0.604 | 0.751 |
| State + History + Resource | 0.838 | 0.679 | 0.820 |
| State + History + Resource + Graph | 0.846 | 0.692 | 0.829 |

The graph-derived features produced a smaller additional improvement:

$$
0.838
\rightarrow
0.846
$$

in accuracy, while macro F1 increased:

$$
0.679
\rightarrow
0.692
$$

The full progression was:

$$
0.666
\rightarrow
0.781
\rightarrow
0.838
\rightarrow
0.846
$$

The largest improvement came from process history.

Resource identity added another substantial increase.

Graph-derived relational context added a smaller but consistent improvement across the reported metrics.

## What This Suggests for CRM Architecture

The original appetite conjecture did not survive the experiments in a useful general form.

Static CRM data performed poorly on future opportunity ranking. Adding interaction summaries helped only modestly. Sequential history became more useful once the target moved closer to the process itself, and the strongest result came from predicting the next transition using current state, process history, resource identity, and local relational context.

The progression in the final experiment was:

$$
0.666
\rightarrow
0.781
\rightarrow
0.838
\rightarrow
0.846
$$

for:

```text
current state
→ state + history
→ state + history + resource
→ state + history + resource + graph context
```

This does not show that every CRM should use a graph database.

It does suggest that the representation used for prediction matters.

### The Relational Representation

A conventional relational CRM might store an opportunity across tables such as:

```text
opportunities
-------------
id
account_id
owner_id
stage
value
created_at
updated_at
```

with related information in other tables:

```text
contacts
activities
users
accounts
documents
stage_history
```

There is nothing inherently incapable about this representation. SQL can reconstruct relationships and history through joins.

For example:

```sql
SELECT
    o.id,
    o.stage,
    a.type,
    u.id AS owner_id,
    h.previous_stage,
    h.changed_at
FROM opportunities o
LEFT JOIN accounts a
    ON a.id = o.account_id
LEFT JOIN users u
    ON u.id = o.owner_id
LEFT JOIN stage_history h
    ON h.opportunity_id = o.id;
```

The issue is not whether a relational database can store the information. It can.

The question is what representation should be exposed to a decision model when the decision depends on a connected process unfolding over time.

A single opportunity row might say:

```text
opportunity_id = 1842
stage          = proposal
owner          = user_17
value          = 120000
```

but the process that produced that row may contain:

```text
created
→ assigned to user_04
→ contacted
→ qualified
→ assigned to user_17
→ proposal created
→ customer replied
→ proposal revised
→ proposal
```

The current row is one projection of that history.

### A Graph Representation

The same process can be represented in terms of entities and their relationships:

```text
Opportunity 1842
├── BELONGS_TO → Account 93
├── OWNED_BY → User 17
├── INVOLVES → Contact 51
├── HAS_EVENT → Call 882
├── HAS_EVENT → Email 901
├── HAS_EVENT → Proposal 44
└── HAS_STATE → Proposal
```

Events can also connect to each other through time:

```text
Call 882
→ FOLLOWED_BY → Email 901
→ FOLLOWED_BY → Proposal 44
→ FOLLOWED_BY → Revision 46
```

and resources can connect multiple cases:

```text
User 17
├── HANDLED → Opportunity 1842
├── HANDLED → Opportunity 2201
└── HANDLED → Opportunity 2390
```

That makes relationships part of the representation itself rather than something reconstructed only when a particular query needs them.

For a prediction at time $t$, the available context can be thought of as a local subgraph:

$$
G_{i,t}
=
(V_{i,t}, E_{i,t})
$$

where:

- $V_{i,t}$ is the set of relevant entities known at time $t$
- $E_{i,t}$ is the set of relationships between those entities

A model can then condition on:

$$
P(
S_{i,t+1}
\mid
S_{i,t},
H_{i,t},
G_{i,t}
)
$$

The experiments do not establish that this is always the best representation. They show that information represented by $H_{i,t}$ and parts of $G_{i,t}$ improved next-state prediction on the process data tested here.

[Appendix A2](#appendix-a2-feature-vectors) and [Appendix A6](#appendix-a6-categorical-and-numerical-preprocessing) give the preprocessing view of this: graph context still has to be converted into model-readable features before a standard classifier can use it.

### Current State Is a Compression

The clearest result from the experiments is that current state can discard information.

Let the complete process history before time $t$ be:

$$
H_{i,t}
=
(
e_{i,0},
e_{i,1},
\ldots,
e_{i,t}
)
$$

and let the CRM derive a current state from that history:

$$
S_{i,t}
=
g(H_{i,t})
$$

Many different histories can produce the same state.

For two cases $a$ and $b$:

$$
H_{a,t} \neq H_{b,t}
$$

while:

$$
g(H_{a,t})
=
g(H_{b,t})
$$

A simple example is:

```text
Case A

new
→ engaged
→ qualified
→ proposal
```

and:

```text
Case B

new
→ engaged
→ qualified
→ engaged
→ qualified
→ proposal
```

Both can be stored as:

```text
stage = proposal
```

but the histories are different.

Whether that difference matters depends on the task.

For terminal-outcome ranking in the real data, the additional information was weak.

For next-state prediction, it mattered considerably more.

That distinction is important. There is no reason to assume that richer process state will improve every prediction problem equally.

## A Stateful CRM

The architecture I find more interesting after these experiments is a CRM where the current record remains useful, but it is treated as one view over an underlying process history.

Conceptually:

```text
entities
    +
events
    +
relationships
    +
time
    ↓
current process state
    ↓
prediction / decision
```

An event might look like:

```json
{
  "type": "proposal_sent",
  "opportunity_id": "1842",
  "actor_id": "user_17",
  "contact_id": "contact_51",
  "timestamp": "2026-08-31T15:42:00Z"
}
```

Another event could record a state transition:

```json
{
  "type": "stage_changed",
  "opportunity_id": "1842",
  "from": "qualified",
  "to": "proposal",
  "actor_id": "user_17",
  "timestamp": "2026-08-31T15:44:00Z"
}
```

The current opportunity record can still be derived and queried normally:

```text
stage = proposal
owner = user_17
```

but the underlying history remains available when a decision depends on how that state was reached.

This also avoids requiring every downstream model to treat the CRM row as a complete description of the business process.

## What Survived

I started these experiments looking for a generalized ranking function for deal appetite.

The synthetic version was recoverable because the latent structure had been placed into the simulation. Real CRM data did not provide evidence for a comparably useful scalar, and increasingly detailed histories did not produce a strong general model of terminal outcomes.

The more consistent result concerned process state.

Static CRM fields were weak for future deal ranking. Real transition history retained modest information within the same process state. When the target was changed to the next observable transition, history produced a much larger improvement, resource identity added more, and graph-derived relational context added a smaller additional gain.

The experiments therefore give me a narrower hypothesis to work with:

> Business state may be better represented for decision-making as a history of connected entities and events than as the current values of those entities alone.

That is a claim about representation, not about Neo4j specifically.

A relational database can store the same underlying information. A graph may simply provide a more direct model when the questions being asked depend on paths, shared entities, transitions, and local context.

The remaining question is whether that advantage survives outside this dataset and whether it is large enough to justify the additional complexity of a graph-native representation.

## Hypothesis Revisited

The original hypotheses were:

### $H_1$

There exists a recoverable ordering function over observable CRM variables:

$$
A_i = f(x_i) + \varepsilon_i
$$

such that $f(x_i)$ preserves useful relative ordering even when $\varepsilon_i$ is unobserved.

### $H_0$

Observable CRM variables do not recover a stable relative ordering of opportunities beyond noise and unobserved variation.

The experiments do not provide enough evidence to support $H_1$ in this general form.

The synthetic experiment recovered the ranking structure that had been explicitly built into the data-generating process. That established that the estimation procedure could recover the systematic component under the assumptions of the simulation, but it did not establish that a comparable one-dimensional ordering function exists in real CRM data.

On real CRM data, static opportunity features produced weak out-of-sample ranking of terminal outcomes. Interaction history and process-history features improved performance in some experiments, but the gains were not strong or consistent enough to establish a stable generalized ordering of opportunities.

I therefore would not reject $H_0$ for the original conjecture.

The later experiments support a narrower result.

Process history can retain information that is absent from the current-state representation, and that information becomes substantially more useful when the target is the next observable process transition rather than a distant terminal outcome.

In the final experiment, next-event accuracy increased from:

$$
0.666
$$

using current state alone, to:

$$
0.781
$$

after adding process history, then to:

$$
0.838
$$

after adding resource identity, and finally to:

$$
0.846
$$

after adding graph-derived relational context.

Those results do not rescue the original $H_1$. They point to a different hypothesis about how business processes should be represented for prediction and decision-making.

## Conclusion

The deal-appetite conjecture began as a search for a generalized opportunity-ranking function. The synthetic experiment showed that such a function can be recovered when the recoverable structure is deliberately built into the data-generating process. The real-data experiments did not show that an analogous scalar exists in ordinary CRM records.

The more durable finding is about representation.

When the target is a distant terminal outcome, richer history is only weakly useful in the datasets tested here. When the target is the next observable process transition, history and relational context matter much more.

That suggests a narrower and more useful direction:

> Business prediction should often be framed around the stateful process that produces the current CRM record, not only around the current record itself.

A graph database is not required for that claim to be true. But graph-shaped representations make the relevant objects easier to name directly: events, actors, accounts, contacts, offers, states, transitions, and the paths between them.

The next question is whether richer graph-native representations can improve decision support in real CRM systems enough to justify their added complexity.

## Appendices

These appendices are numbered by topic. Appendix A covers notation, feature construction, and preprocessing. Appendix B covers evaluation metrics. Appendix C covers data and reproducibility.

### Appendix A1: Libraries and Notation

This appendix explains the notation and library mechanics used throughout the experiments. It is meant to make the article readable even if the reader has only seen introductory statistics or discrete math.

The experiments were written in Python and run in Jupyter.

The main libraries were:

- **pandas** for loading, joining, filtering, grouping, and transforming CRM and event-log data
- **NumPy** for random-number generation and vectorized numerical operations in the synthetic experiments
- **scikit-learn** for preprocessing, model fitting, and evaluation
- **SciPy** for Spearman rank correlation
- **Matplotlib** for plots and visual inspection of results

**Notation guide**

The notation is compact because the same structure appears many times.

| Symbol | Plain meaning |
|---|---|
| $i$ | the index of one opportunity, application, case, or event snapshot |
| $j$ | another index, usually used when comparing two examples |
| $n$ | the number of examples |
| $x$ | one input value |
| $\mathbf{x}_i$ | the full feature vector for example $i$ |
| $X$ | the full feature table or matrix |
| $y_i$ | the observed label for example $i$ |
| $\hat{p}_i$ | the model's estimated probability for example $i$ |
| $\hat{y}_i$ | the model's predicted class for example $i$ |
| $S_i$ | systematic appetite in the synthetic experiment |
| $A_i$ | total latent appetite in the synthetic experiment |
| $\epsilon_i$ | unobserved variation for example $i$ |
| $\rho_s$ | Spearman rank correlation |

A subscript tells us **which example** we are talking about.

For example:

$$
\mathbf{x}_7
$$

means "the feature vector for example 7."

A superscript in this article is usually a **label**, not a power.

For example:

$$
\mathbf{x}_i^{(s)}
$$

means the static CRM feature vector for opportunity $i$.

It does not mean $\mathbf{x}_i$ multiplied by itself.

The hat symbol means **estimated by the model**.

For example:

$$
p_i
$$

means the true probability in the synthetic experiment, while:

$$
\hat{p}_i
$$

means the model's estimated probability.

The vertical bar means "given" or "conditional on."

For example:

$$
P(Y_i = 1 \mid \mathbf{x}_i)
$$

means:

> the probability that example $i$ has outcome $1$, given the features in $\mathbf{x}_i$.

The membership symbol means "is an element of."

For example:

$$
\sigma \in [0,5]
$$

means:

> $\sigma$ is chosen from the interval between $0$ and $5$.

In the noise-sensitivity experiment, this was implemented as a finite grid:

```python
noise_levels = np.arange(0, 5.25, 0.25)
```

So the computer did not test every real number between $0$ and $5$. It tested:

```text
0.00, 0.25, 0.50, ..., 5.00
```

The normal-distribution notation:

$$
x_i \sim \mathcal{N}(0,1)
$$

means:

> draw $x_i$ from a normal distribution with mean $0$ and standard deviation $1$.

### Appendix A2: Feature Vectors

A model cannot read a CRM row the way a person reads a spreadsheet. It needs a list of values.

For one opportunity $i$, that list is called a feature vector:

$$
\mathbf{x}_i
$$

The bold $\mathbf{x}$ means that the input contains several values, not just one value.

Suppose one opportunity has:

```text
product          = GTX Pro
sector           = Technology
sales_price      = 82,000
employees        = 220
is_subsidiary    = 0
engage_month     = 3
```

Before a standard model can use that row, text values must be converted into numbers. After preprocessing, the model might receive this vector:

```text
feature name          value
-------------------   -----
product_GTX_Basic       0
product_GTX_Pro         1
product_MG_Special      0
sector_Finance          0
sector_Technology       1
scaled_sales_price      0.42
scaled_employees       -0.18
is_subsidiary           0
engage_month            3
```

The same thing can be written as a mathematical vector:

$$
\mathbf{x}_i
=
\begin{bmatrix}
0 \\
1 \\
0 \\
0 \\
1 \\
0.42 \\
-0.18 \\
0 \\
3
\end{bmatrix}
$$

Each row of the vector is one input feature.

When there are many opportunities, the model receives a table of feature vectors:

$$
X
=
\begin{bmatrix}
--- \mathbf{x}_1 --- \\
--- \mathbf{x}_2 --- \\
\cdots \\
--- \mathbf{x}_n ---
\end{bmatrix}
$$

Each row is one example. Each column is one feature.

A small feature table might look like this:

| opportunity | product_GTX_Pro | sector_Technology | scaled_sales_price | won |
|---|---:|---:|---:|---:|
| 1 | 1 | 1 | 0.42 | 1 |
| 2 | 0 | 1 | -0.60 | 0 |
| 3 | 1 | 0 | 1.10 | 1 |

The model is trained on the feature columns:

$$
X
$$

and the label column:

$$
y
$$

For a binary sales model:

```text
y_i = 1  if the deal was won
y_i = 0  if the deal was lost
```

So the model learns a relationship of the form:

$$
\mathbf{x}_i
\rightarrow
y_i
$$

or, for probability prediction:

$$
\mathbf{x}_i
\rightarrow
\hat{p}_i
$$

where $\hat{p}_i$ is the model's estimated probability.

### Appendix A3: Feature Groups

The article uses superscripts to label different groups of inputs.

For example:

$$
\mathbf{x}_i^{(s)}
$$

means static CRM features, such as product, sector, salesperson, and deal value.

$$
\mathbf{x}_i^{(b)}
$$

means behavioral or interaction-derived features, such as interaction count, reply timing, or channel mix.

The combined model can therefore be read as:

$$
P
\left(
Y_i = 1
\mid
\mathbf{x}_i^{(s)},
\mathbf{x}_i^{(b)}
\right)
$$

which means:

> the probability that opportunity $i$ is won, given its static CRM features and behavioral features.

### Appendix A4: Logistic Regression

In a binary model, logistic regression first converts the feature vector into a single score:

$$
z_i
=
\beta_0
+
\beta_1x_{i1}
+
\beta_2x_{i2}
+
\cdots
+
\beta_kx_{ik}
$$

Here:

- $x_{i1}$ is the first feature for example $i$
- $x_{i2}$ is the second feature for example $i$
- $k$ is the number of features
- $\beta_0$ is the intercept
- $\beta_1,\ldots,\beta_k$ are learned coefficients

A tiny worked example:

```text
features:
    product_GTX_Pro      = 1
    sector_Technology    = 1
    scaled_sales_price   = 0.42

learned coefficients:
    intercept            = -0.30
    product_GTX_Pro      = 0.50
    sector_Technology    = 0.20
    scaled_sales_price   = 0.80
```

The model score is:

$$
z_i
=
-0.30
+
0.50(1)
+
0.20(1)
+
0.80(0.42)
=
0.736
$$

That score is not yet a probability. Logistic regression applies the logistic function:

$$
\hat{p}_i
=
\frac{1}{1 + e^{-z_i}}
$$

So:

$$
\hat{p}_i
=
\frac{1}{1 + e^{-0.736}}
\approx
0.676
$$

The model would estimate about a $67.6\%$ probability of the positive class.

In code, scikit-learn does this after the model has been fit:

```python
predicted_probability = model.predict_proba(X_test)[:, 1]
```

The `[:, 1]` means:

> take every row, and take column 1 of the probability output.

For a binary classifier, column 0 is usually the probability of class `0`, and column 1 is the probability of class `1`.

### Appendix A5: Histories as Feature Vectors

For process-history models, the raw object is no longer just one CRM row. It is a history:

```text
A_Submitted
-> A_Concept
-> A_Accepted
-> O_Create Offer
```

To use a standard classifier, that history still has to be converted into features.

For example:

```text
current_activity        = A_Accepted
previous_activity       = A_Concept
event_position          = 3
elapsed_hours           = 12.7
unique_activity_count   = 3
offer_event_count       = 0
workflow_event_count    = 1
```

Those values form the feature vector for that prediction time.

For next-event prediction, each training row contains information available at time $t$, and the label is the event that happened at time $t+1$.

| case | event position $t$ | current activity | previous activity | elapsed hours | label: next activity |
|---|---:|---|---|---:|---|
| C-001 | 3 | A_Accepted | A_Concept | 12.7 | O_Create Offer |
| C-002 | 5 | W_Complete application | W_Handle leads | 3.4 | A_Complete |

This is the same supervised-learning structure as before:

$$
\text{known inputs at time } t
\rightarrow
\text{observed outcome at time } t+1
$$

The important rule is that features must be computed only from information available at or before time $t$.

### Appendix A6: Categorical and Numerical Preprocessing

For static CRM models, categorical and numerical fields require different preprocessing.

Categorical fields are fields with names or categories, such as:

```text
product = GTX Pro
sector  = Technology
office  = West
```

They were converted into numeric columns using `OneHotEncoder`:

```python
from sklearn.preprocessing import OneHotEncoder
```

For example:

```text
product = GTX Pro
```

may become:

```text
GTX Basic   GTX Pro   MG Special
    0          1           0
```

The setting:

```python
OneHotEncoder(handle_unknown="ignore")
```

means:

> if a category appears in the test set that was not seen during training, do not crash; encode it as all zeros for that feature group.

Numerical fields are already numbers, such as revenue, employee count, elapsed time, or event count.

They were standardized using `StandardScaler`:

```python
from sklearn.preprocessing import StandardScaler
```

Standardization converts a raw value $x$ into:

$$
z
=
\frac{x - \mu}{\sigma}
$$

where $\mu$ is the mean of that feature in the training data and $\sigma$ is its standard deviation.

For example, if the average deal value in the training set is $50{,}000$ and the standard deviation is $20{,}000$, then a deal worth $70{,}000$ becomes:

$$
z
=
\frac{70{,}000 - 50{,}000}{20{,}000}
=
1
$$

That means the deal is one training-set standard deviation above the training-set average.

A deal worth $30{,}000$ would become:

$$
z
=
\frac{30{,}000 - 50{,}000}{20{,}000}
=
-1
$$

That means the deal is one training-set standard deviation below the training-set average.

The two preprocessing steps were combined using `ColumnTransformer`:

```python
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler

preprocessor = ColumnTransformer(
    transformers=[
        (
            "categorical",
            OneHotEncoder(handle_unknown="ignore"),
            categorical_features,
        ),
        (
            "numeric",
            StandardScaler(),
            numeric_features,
        ),
    ]
)
```

This says:

```text
apply OneHotEncoder to categorical_features
apply StandardScaler to numeric_features
combine the resulting columns into one model-readable matrix
```

### Appendix A7: Training-Only Preprocessing

Most models used scikit-learn pipelines so that preprocessing was fit only on the training data and then applied unchanged to the temporal holdout.

This matters because preprocessing itself can leak future information.

Suppose the model is evaluated on later opportunities, but the scaler is fit using both earlier and later opportunities. Then the training process has already seen information about the distribution of the future test set.

That gives the model information it would not have had at the time of prediction.

The safer workflow is:

```text
split rows into train and test by time
fit preprocessing on training rows only
transform training rows
train model
transform future test rows using the same fitted preprocessing
evaluate model
```

The preprocessing step and classifier were wrapped in a `Pipeline`:

```python
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression

pipeline = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=2000)),
    ]
)
```

Training and prediction then followed the standard scikit-learn pattern:

```python
pipeline.fit(X_train, y_train)

predicted_probability = pipeline.predict_proba(X_test)[:, 1]
```

Inside `pipeline.fit`, scikit-learn does this:

```text
preprocessor.fit_transform(X_train)
classifier.fit(transformed_X_train, y_train)
```

Inside `pipeline.predict` or `pipeline.predict_proba`, it does this:

```text
preprocessor.transform(X_test)
classifier.predict(...) or classifier.predict_proba(...)
```

Notice that the test data is transformed, but the preprocessing rules are not refit on the test data.

For the next-event prediction experiment, the target was multiclass rather than binary. Those models used an `SGDClassifier` with logistic loss:

```python
from sklearn.linear_model import SGDClassifier

classifier = SGDClassifier(
    loss="log_loss",
    max_iter=1000,
    tol=1e-3,
    random_state=42,
)
```

In this setting, the classifier is not asking "won or lost?" It is asking:

```text
which activity is most likely to happen next?
```

So the prediction is a class label:

```python
prediction = pipeline.predict(X_test)
```

and the evaluation uses multiclass metrics.

### Appendix B1: Evaluation Tasks

The experiments use two kinds of prediction tasks.

The first kind is **binary ranking**: rank won deals above lost deals, or successful applications above unsuccessful applications.

The second kind is **multiclass classification**: predict exactly which event happens next.

Those tasks need different metrics.

### Appendix B2: ROC AUC

For binary opportunity-outcome experiments, ranking performance was measured with ROC AUC:

```python
from sklearn.metrics import roc_auc_score

auc = roc_auc_score(
    y_test,
    predicted_probability
)
```

ROC AUC measures how often the model ranks a randomly selected positive case above a randomly selected negative case.

Here is a small worked example.

| deal | observed outcome | model score |
|---|---:|---:|
| A | 1 | 0.90 |
| B | 0 | 0.80 |
| C | 1 | 0.70 |
| D | 0 | 0.20 |

There are two positive examples, A and C, and two negative examples, B and D.

That creates four positive-negative pairs:

| pair | correct ranking? | reason |
|---|---|---|
| A vs B | yes | $0.90 > 0.80$ |
| A vs D | yes | $0.90 > 0.20$ |
| C vs B | no | $0.70 < 0.80$ |
| C vs D | yes | $0.70 > 0.20$ |

Three of the four pairs are ordered correctly, so:

$$
\operatorname{AUC}
=
\frac{3}{4}
=
0.75
$$

With ties, AUC gives half credit. The general counting version is:

$$
\operatorname{AUC}
=
\frac{
\#\text{correct pairs}
+
0.5 \cdot \#\text{tied pairs}
}{
\#\text{positive examples}
\cdot
\#\text{negative examples}
}
$$

This is why AUC is a ranking metric. It does not require the predicted probability to be perfectly calibrated. It asks whether positive examples tend to receive higher scores than negative examples.

An AUC of:

$$
0.5
$$

is equivalent to random ranking, while:

$$
1.0
$$

represents perfect ranking.

In the synthetic experiment, AUC was useful because the observed outcome $Y_i$ was stochastic. Even if a deal had high true probability, it could still fail to close. AUC asks whether the model tends to rank the closed deals higher overall.

### Appendix B3: Spearman Rank Correlation

The synthetic experiment also compares the model's predicted ranking against latent quantities that are known only because the data was simulated.

This was measured using Spearman rank correlation:

```python
from scipy.stats import spearmanr

rank_corr, _ = spearmanr(
    predicted_probability,
    test_appetite
)
```

Spearman correlation first converts values into ranks.

For example:

```text
raw scores:   0.20   0.90   0.40
ranks:          1      3      2
```

Then it compares the rank ordering of one variable with the rank ordering of another variable.

Here is a worked example with four opportunities.

| opportunity | model score $\hat{p}$ | rank of $\hat{p}$ | latent appetite $A$ | rank of $A$ | rank difference $d$ |
|---|---:|---:|---:|---:|---:|
| A | 0.10 | 1 | 5.0 | 1 | 0 |
| B | 0.85 | 4 | 8.0 | 3 | 1 |
| C | 0.40 | 2 | 6.0 | 2 | 0 |
| D | 0.70 | 3 | 10.0 | 4 | -1 |

Without ties, Spearman correlation can be computed as:

$$
\rho_s
=
1
-
\frac{
6\sum_i d_i^2
}{
n(n^2 - 1)
}
$$

In this example:

$$
\sum_i d_i^2
=
0^2 + 1^2 + 0^2 + (-1)^2
=
2
$$

and:

$$
\rho_s
=
1
-
\frac{6(2)}{4(4^2 - 1)}
=
1
-
\frac{12}{60}
=
0.8
$$

A value near:

$$
1
$$

means the two rankings are almost the same.

A value near:

$$
0
$$

means there is no consistent ranking relationship.

A value near:

$$
-1
$$

means the rankings are almost opposite.

This is why the synthetic experiment can distinguish:

$$
\rho_s(\hat{p}, A)
$$

from:

$$
\rho_s(\hat{p}, S)
$$

The first compares the model ranking with total latent appetite. The second compares it with the systematic ordering recoverable from observable features.

This distinction matters because:

$$
A_i = S_i + \epsilon_i
$$

The model can learn structure in $S_i$ because $S_i$ is built from observable inputs. It cannot directly recover the realized value of $\epsilon_i$ because that term is hidden.

### Appendix B4: Accuracy

For next-event prediction, the target has many possible classes, so the experiment reports accuracy, macro F1, and weighted F1.

Accuracy measures the proportion of predictions that are exactly correct.

For example, if the model makes $100$ predictions and gets $84$ right:

$$
\operatorname{Accuracy}
=
\frac{84}{100}
=
0.84
$$

In scikit-learn:

```python
from sklearn.metrics import accuracy_score

accuracy = accuracy_score(
    y_test,
    prediction
)
```

Accuracy is easy to understand, but it can hide poor performance on rare classes.

Suppose a dataset has $100$ next events:

```text
80 are A_Submitted
15 are O_Create Offer
5 are A_Cancelled
```

A model that mostly predicts the common class can look decent by accuracy while still being bad at rare but important events.

### Appendix B5: Precision, Recall, and F1

F1 is built from precision and recall.

For one class, precision asks:

> Of the cases the model predicted as this class, how many were actually this class?

Recall asks:

> Of the cases that truly belonged to this class, how many did the model find?

For one class:

$$
\operatorname{Precision}
=
\frac{
\operatorname{true\ positives}
}{
\operatorname{true\ positives}
+
\operatorname{false\ positives}
}
$$

and:

$$
\operatorname{Recall}
=
\frac{
\operatorname{true\ positives}
}{
\operatorname{true\ positives}
+
\operatorname{false\ negatives}
}
$$

Then:

$$
F_1
=
2
\cdot
\frac{
\operatorname{Precision}
\cdot
\operatorname{Recall}
}{
\operatorname{Precision}
+
\operatorname{Recall}
}
$$

A small example for one class:

```text
true positives   = 8
false positives  = 2
false negatives  = 4
```

Then:

$$
\operatorname{Precision}
=
\frac{8}{8 + 2}
=
0.80
$$

and:

$$
\operatorname{Recall}
=
\frac{8}{8 + 4}
\approx
0.667
$$

so:

$$
F_1
=
2
\cdot
\frac{0.80 \cdot 0.667}{0.80 + 0.667}
\approx
0.727
$$

In scikit-learn:

```python
from sklearn.metrics import f1_score

macro_f1 = f1_score(
    y_test,
    prediction,
    average="macro"
)

weighted_f1 = f1_score(
    y_test,
    prediction,
    average="weighted"
)
```

### Appendix B6: Macro F1 and Weighted F1

Macro F1 calculates F1 separately for each class and gives every class equal weight.

Weighted F1 also calculates F1 separately, but weights classes according to how often they occur.

Suppose a next-event model has these class-level F1 scores:

| next-event class | support | class F1 |
|---|---:|---:|
| A_Submitted | 80 | 0.90 |
| O_Create Offer | 15 | 0.50 |
| A_Cancelled | 5 | 0.10 |

Macro F1 is:

$$
\operatorname{Macro\ F1}
=
\frac{0.90 + 0.50 + 0.10}{3}
=
0.50
$$

Weighted F1 is:

$$
\operatorname{Weighted\ F1}
=
\frac{
80(0.90)
+
15(0.50)
+
5(0.10)
}{
100
}
=
0.80
$$

The same model can therefore have high weighted F1 and much lower macro F1.

That pattern means:

> the model is doing much better on common classes than rare classes.

The increase in macro F1:

$$
0.604
\rightarrow
0.679
$$

therefore matters because it suggests that resource context improved prediction across the set of next-event classes, not only for the most frequent transition.

### Appendix C1: Data and Reproducibility

The notebooks are published in the companion GitHub repository:

```text
https://github.com/vbnovikov/deal-appetite-experiment
```

Raw datasets are not included in the repository. They should be downloaded from their original sources and placed under the local `data/` directory described in `DATA_SOURCES.md`.

The repository is licensed under the MIT License. The datasets remain subject to their original source licenses.
