Why Data and LLMs Alone Cannot Create Scientific Outcomes

Enterprises have valuable scientific data and increasingly capable models. A scientific harness and enterprise transformation turn that capability into work R&D can trust and use.

I hear a familiar opinion in enterprise AI discussions: connect a frontier LLM to proprietary data and domain value will follow. It feels intuitively right, but it is incomplete. Why? Don't we just need intelligence plus data to solve a scientific task? The answer is no. Both are necessary, but not sufficient.

In R&D, an outcome is completed scientific work that can withstand expert review and enter a decision process. A patient responder analysis, for instance, must test competing explanations behind a signal, find supporting and contradictory evidence, preserve provenance, make uncertainty visible, and produce an analysis that experts can inspect. Intelligence can reason over the problem, and data can supply evidence. But, neither defines how the work should be carried out, checked, reviewed, or embedded in a decision. That is why the equation needs two more terms:

Scientific outcome = LLM + data + scientific harness + enterprise transformation.

The LLM contributes general intelligence, reasoning, coding, and tool use. Data contributes evidence and organizational context. The scientific harness encodes methods, retrieval, memory, evaluation, uncertainty, and human checkpoints. Enterprise transformation embeds that work in roles, review standards, decision rights, and learning processes.

Figure 1. LLM and data become a scientific outcome only when the scientific harness and enterprise transformation complete the system.

Coding agents show us why the harness matters

Pharma IT teams are trained to compare capabilities: model quality, data access, connectors, agents, security, governance, and cost. All of those matter. Still, the checklist cannot show whether the assembled system will complete the scientific work. Software development gives us an established benchmark for seeing what the harness changes.

Task-specific benchmarks make the harness effect visible. Terminal-Bench 2.0 held Claude Opus 4.5 and the task set constant while changing the agent around it. Terminus 2, an agent harness, resolved 57.8% of tasks, compared with 52.1% for Claude Code. That is a 5.7 percentage-point difference using the same model! The paper also reports 3.9 million input tokens for Terminus 2, versus 256.9 million for Claude Code. Terminus 2 completed more of the tested work while processing far less input context.

The result is specific to terminal work, not scientific work. The point is that the harness determines how efficiently and reliably a model's capability is expressed for a defined task. Scientific work needs its own task-specific methods, evidence controls, evaluation, and human checkpoints.

Science raises the bar further. Coding often has fast executable tests. Scientific work must retrieve incomplete evidence, preserve provenance, distinguish association from causality, manage uncertainty, and support scientific interpretation that experts can challenge. A scientific harness must therefore encode a richer method and judgment loop.

Figure 2. With Claude Opus 4.5 and the benchmark fixed, Terminus 2 completed more tasks while using far fewer reported tokens than Claude Code or OpenHands.

The tools available to the harness: data must become scientifically findable

A harness orchestrates model intelligence, but it also needs tools. For instance, the Causaly Agent has access to two knowledge graphs and a scientific search system. Even after an enterprise connects its data, the agent still needs reliable ways to find the right evidence. That is straightforward in a file system and far harder in an unbiased scientific search.

Enterprise data spans external literature, trials, patents, and databases, as well as internal reports, experiments, decisions, and negative findings. Possessing data is different from making it scientifically findable. Scientific retrieval and knowledge graphs normalize entities, connect relationships, rank evidence, and preserve provenance. Otherwise, the model sees what happened to be retrieved rather than the evidence the question requires.

Connecting data is necessary, but not enough – it must be findable. Think of a brilliant mathematician limited to pen and paper versus a good one with a computer (a tool). Capability matters. So does the system through which it is applied.

Enterprise transformation makes scientific work repeatable

A strong harness, the right data, and model intelligence can create a personal productivity tool. Workflow transformation at the department level does not follow from individual productivity alone. Review cadence, handoffs, work product definitions, incentives, decision rights, correction paths, and reuse norms have to move with the technology.

That requires a transformation program: bottom-up training and champion communities, top-down governance, and process or incentive redesign where needed. Without that layer, even strong individual results remain isolated. With it, completed scientific work becomes repeatable across a department.

Figure 3. Individual productivity becomes enterprise value when the operating model moves with the technology.
Value comes from the complete system

Life sciences enterprises possess unusually valuable data. Harvesting that value is not just a matter of placing an LLM on top. Terminal-Bench shows how much the harness can change execution before the system even faces scientific evidence and expert judgment.

Decision makers should resist the shortcut that data plus an LLM equals value. Scientific outcome = LLM + data + scientific harness + enterprise transformation. All four are required.

Get started with Causaly

Ready to transform the way your R&D teams discover and deliver? Take the first step - see Causaly for yourself.

Request a demo