Client story · Northline · 2025
Qualified conversations, not longer forms
Northline’s demo form was collecting a number sales never used. A gradient-boosted model, a form heatmap, and exit interviews put the leave on one field. Taking it off the page raised qualified leads 38%.
+38%qualified leads
- Capture
- Model
- Insight
- Interface
- Test
- Hold
Summary
Northline sells workflow software to operations teams. The demo form asked for role, problem, a work email, and an annual budget. Sales used the first three. The fourth was an empty column in the CRM that had been promoted into a required pixel. People who could not price a budget yet left. People who guessed were often the wrong conversation.
We joined every field focus, every pause over two seconds, and every abandon to the later call outcome. A gradient-boosted model ranked which of those features predicted a lead sales would mark qualified. Time on the budget field was the strongest negative feature. The heatmap agreed: a bright band across that one row, warm on role and problem, cooler on email. Exit interviews said the same thing without the model. They did not have the number yet.
The landing that shipped kept role, problem, and work email. Budget moved to the call agenda. Over six weeks, split by session, qualified leads rose 38%. The share of submits that became a held call rose 21%. A second cohort repeated the lift. Raw form completes were reported and were not allowed to decide.
The question the form was pretending to answer
Ira Sen, who runs revenue at Northline, described the history plainly. Fields had been added because the CRM had empty columns. The form got longer. The calls did not get better. That sentence was the brief. It is also a warning about outcome labels. If the recorded success is “form submitted,” a longer form can look like a better form right up until sales stops returning the calls.
The decision we were willing to be wrong about was narrow. Does asking for an annual budget on the first screen improve the quality of the conversation, or does it end conversations that would have been worth having? Quality had a definition before any model was fit. A qualified lead was a submit that sales later marked as a real fit. A held call was a submit that became a meeting on the calendar. Both labels lived downstream of the form, which is the only reason they were allowed to judge it.
What was recorded, and what was refused
The event stream was ordinary and strict. For each session we kept the order of field focus, the time spent in each field, pauses longer than two seconds, whether the session abandoned before submit, the device class, and the campaign only as a covariate we did not let drive the design. After the call, the sales disposition was joined back on the session that produced the form.
The join is where this kind of work usually cheats. A model that sees the call outcome at scoring time is not predicting anything. It is reading the label. Training used completed histories only. The scores were used to explain which features separated qualified conversations from leaves. The interface change was then tested forward, on sessions that had not been used to build the explanation. Campaign, company size, and the page they came from were held as controls. They were not allowed to become the story if the field-level features already carried it.
Cleaning the record before anyone ranked a feature
Sessions with no field focus were dropped. They are bounces, and a bounce is not evidence about a question. A pause under two tenths of a second on budget was coded as a skip, not a read. Counting those as “considered and rejected” would have inflated the field’s importance in the wrong direction. Submits with no later sales disposition were held out of the qualified label and out of the negative label. An unlabeled call is not a bad lead. It is a missing label, and missing labels were not filled in.
Duplicate submits from the same work email inside an hour were collapsed to the first completed session. The later one was usually a correction, and treating it as a second person would have double-counted both the dwell and the outcome. Internal addresses and the sales team’s own test rows were removed before the model saw the table. None of this is glamorous. It is the difference between a feature that means “people struggle here” and a feature that means “our logging is noisy.”
A gradient-boosted model pointed at the conversation
The model was a gradient-boosted classifier. The target was the sales disposition, not the submit button. Features were the field-level timings, the pause counts, the order of focus, whether budget was touched at all, and whether the session ended before submit. We did not hand the model the text people typed. The question was about the ask, not about the prose.
Budget dwell dominated the negative class. Role and problem were weak and slightly positive: time there looked like someone describing a real job. Email was nearly neutral, which matched the heatmap. The middle band of the form, the budget row, was the hottest zone. Above it the cells were warm. Below it they cooled. A model and a picture that agree are still not a result. They are a hypothesis with a mechanism. The mechanism was: people were not refusing the product. They were refusing a number they did not have yet.
Exit interviews were the check on that sentence. People who had abandoned were asked what they were doing in the last field they touched. The ones who stopped on budget did not describe a pricing objection. They described not knowing the figure, and not wanting to invent one in front of a vendor. That is a different problem from “the product is too expensive,” and it asks for a different pixel.
The landing, drawn on the grid
The shipped screen is a landing and nothing else. A title, three fields, one call to action: Book a conversation. Role, problem, work email. The budget row is still on the drawing, dashed, because that is where the tension was. It is not on the page a visitor fills in. Budget is a line on the call agenda, where a person can say they do not know yet.
Before, four fields stood in a column and the button waited under the number. After, the button follows the email. Nothing was added to compensate. No progress bar, no “typical budget” hint, no second step that re-asked the same question in softer words. The research had named one field. The interface was allowed to remove one field.
- Quiet
- Low
- Warm
- Hot
- Tension
Landing blueprint for Northline. Title, three fields, a call-to-action, and a dashed budget field carrying the tension.
Where attention actually sat
The heatmap is an 8 by 8 reading of the old form, not a decoration. The hottest band is the budget field. Role and problem above it are warm. The email field below is cooler. Read next to the model, the picture stops being “people look at the middle of forms.” It becomes “the middle of this form is the question that ends the session.”
We kept the heatmap in the record because a feature importance chart is easy to over-read. A bright cell on a real layout is harder to argue with, and it told the same story as the interviews. Three instruments, one field.
What was said in the room
The design was not a surprise to the people who would have to live with it. Ira’s worry was that sales would ask for the number on the call. That was accepted. A conversation that includes “we don’t have a budget yet” is a better conversation than a form that loses the person before the call exists. Leah’s condition was the one that governed the test: if more people submit and fewer are worth a call, the short form does not ship.
Six weeks, and a second cohort that had to agree
Half of new sessions saw the short form for six weeks. Assignment was by session, not by account, because most of these people had no account yet. The primary measure was leads sales marked qualified. The hold was the share of submits that became a meeting. Submit count was watched so we would notice a hollow win, and it was not the number with a veto except in one direction: a rise in submits with a fall in qualified calls would have killed the change.
Qualified leads rose 38%. Form-to-call conversion rose 21%. A second cohort, held out from the first read, repeated the direction of both. That repetition is what “hold” means here. A single split can flatter a week of traffic. Two splits that agree are a change we were willing to leave on the page.
The pipeline only worked because the label on the outcome was the sales conversation, not the submit button. Point the same model at form completes and the budget field looks helpful. It filters. Point it at qualified conversations and the filter is the leak. The advanced part of the work was not the booster. It was refusing the convenient label.
| Control | Form with the budget field |
|---|---|
| Variant | Form without it |
| Sample | 6 weeks, split by session |
| Primary | +38% leads sales marked qualified |
| Held | +21% of submits became a held call |