Advertisement

Eddie AI Trains Editing Model Small Enough to Run Locally – Claims 47% Win Rate Against Kimi K3

August 20th, 2026Jump to Comment Section2
Eddie AI Trains Editing Model Small Enough to Run Locally – Claims 47% Win Rate Against Kimi K3

Eddie AI has published early research that points somewhere more interesting than another feature update. According to the company, it took a nine-billion-parameter open model and trained it to find a story in raw footage using a larger frontier model as its teacher. Eddie AI says the smaller model was then preferred in 47% of head-to-head comparisons with Kimi K3. What is less clear from the material provided is exactly how those comparisons were conducted. Still, the underlying idea is worth watching: if a relatively small model can handle meaningful editorial decisions, future AI editing tools may not necessarily depend on large cloud-based models.

The research used two specific models. Kimi K3, the open-weight frontier model Moonshot AI released on July 16, 2026, at 2.8 trillion parameters, generated the teacher edits. Qwen3.5-9B, part of Alibaba’s smaller Qwen model family, served as the student. The task was deliberately narrow: finding a compelling story in raw footage, working at the transcript level and evaluating the results on a held-out set of roughly 100 real projects with real prompts. Out of the box, the 9B model was preferred over Kimi on around 10% of edits. Trained on Kimi’s reasoning traces using LoRA, that rose to roughly 32%. After the team changed which trajectories the student learned from, it peaked at 47%.

What a 47% win rate does and does not mean

A 50% win rate would suggest approximate parity within this particular comparison and held-out set. It would not mean parity across video editing in general. Eddie AI makes that limitation clear and describes the 47% result as preliminary.

The size gap is what makes the result interesting. At 2.8 trillion parameters, Kimi K3 is more than 300 times larger than the 9B student model. Coming close to parity on one narrowly defined editorial task with a much smaller model is therefore more interesting than the percentage alone might suggest. But it remains a result tied to a specific task, dataset, and training approach rather than evidence that the smaller model can match Kimi more broadly.

Eddie AI integrates with Iconik – Image credit: Eddie AI

Short social videos were the test case

Eddie AI chose short-form social edits as the test case because they impose tight editorial constraints. The cut has to hook a viewer quickly, fit within a limited duration, and still hold together as a story. Weak soundbite selection becomes difficult to hide.

For an editor, a model that performs better at that task could potentially produce a stronger first assembly and reduce the time spent searching through footage for the narrative spine. It is the same broader problem Eddie AI has been addressing through its Scripted mode and contextual B-roll placement releases, but approached here from the model side rather than the interface side.

More data was not the answer

The most useful finding in the study has little to do with the headline percentage. Eddie AI reports that simply generating more teacher data produced limited gains. Larger improvements came from how the training trajectories were selected and refined.

The process ran in three steps. First, the team generated reasoning traces from the frontier teacher. It then used best-of-N sampling and filtering to retain the stronger examples, meaning the model produced several candidate outputs and the pipeline selected among them rather than training on the first result. Finally, the student generated its own reasoning traces and was retrained iteratively on its strongest outputs.

One side effect is worth noting. At its strongest checkpoint, the student’s reasoning traces were also about 30% shorter than before. In practice, shorter reasoning can reduce inference cost, which adds to the deployment advantage of using a much smaller model.

Image credit: Eddie AI

How reliable is the judge?

Comparisons like this are scored by an automated judge, so the judge itself also needs checking. Eddie AI ran Kimi against Kimi and found that scores varied by only about ±2% across runs, suggesting the evaluation was reasonably stable. The company also asked an in-house video professional to independently rank 10 pairs; the automated judge agreed with the human on seven, while two of the three disagreements fell into the editor’s own “not sure” category.

That is a useful sanity check, but not a validation. Ten pairs is a very small sample, and model-based judging can introduce its own biases into this kind of evaluation. The result is best treated as a directional signal rather than a definitive benchmark. It is not peer-reviewed, and Eddie AI does not present it as one.

The size of the model is the actual story

Eddie AI says the student remained small enough to run on a single, relatively accessible GPU rather than requiring a large inference cluster. A 9B model quantized to 4-bit can fit within the memory range of many workstation GPUs, although actual requirements depend on the inference setup, context length, and other overhead.

That is where the research becomes more relevant to professional post. A model capable of running on-premises could keep footage inside an organization rather than sending it to a cloud service, which matters for broadcasters, sports rights holders, agencies working under NDA, and productions with restrictions on cloud-assisted tooling. It could also make it possible to build more organization-specific editorial workflows without sending that material or associated data to an external provider.

Feedback Mode v2 – Image credit: Eddie AI

The control argument

That last point is the one Eddie AI is leaning on hardest. “Customers are not opposed to AI training; they are opposed to losing control of the knowledge they create,” said Shamir Allibhai, co-founder and CEO of Eddie AI. “Our goal is to give every customer a system that learns only with their permission, improves exclusively for them, and keeps their creative advantage theirs.”

The company sets out four conditions for training it considers acceptable: it should be transparent, happen only with explicit permission, remain under the customer’s control, and make that customer the sole beneficiary of the resulting improvements. That makes Eddie AI’s approach as much a statement about how it wants customer-specific training to work as it is a technical roadmap.

Alexander Terekhov, co-founder and head of AI at Eddie AI, put the technical case more plainly. “A general-purpose model has to be good at everything. A specialized model only has to become exceptionally good at editing video,” he said. “This result gives us confidence that smaller, customer-specific systems can close the gap on valuable editorial decisions while remaining practical to deploy.”

Image credit: Eddie AI

What has not been announced

Quite a lot, and it is worth being clear about it. There is no product here, no release date, no pricing, and no announced on-premises tier. Eddie AI describes the longer-term goal as a model that understands media, learns a filmmaker’s editorial language, remembers how a customer selects B-roll, preserves visual continuity, and eventually handles music and motion graphics. This experiment, however, covers transcript-level story construction only. The company also says frontier models remain central to Eddie AI as the product exists today.

For more context on the current product, see our interview with Shamir Allibhai and our coverage of the v3 Night Shift workflow.

Eddie AI at IBC in Amsterdam

If you are heading to IBC 2026 in September and want to meet the team behind Eddie AI, make sure to visit them at their booth.

Price and availability

Eddie AI is available now for Mac and Windows, with a browser version also available. Users can try it for free. Full details are on Eddie AI’s pricing page, while the app and downloads are available on the Eddie AI website.

Would a model that learns your house style and never leaves your machine change how you feel about AI in the edit suite, or is the whole idea a non-starter regardless of where it runs? Let us know in the comments below!

2 Comments

Filter:
all
Sort by:
latest
Filter:
all
Sort by:
latest