Skip to content

Fabryka AI / Our vision

More intelligence
per GPU.

AI infrastructure should do more than serve models. It should help us discover which models work, make them more efficient, and turn those improvements into better products.

We’re building Fabryka as a connected system for evaluating, routing, adapting, and serving open models—with Polish language and workloads at its center.

The API is the product today.
The factory behind it is the long-term advantage.

01 / A connected system

One factory.
Every part makes the others better.

A benchmark tells us where a model succeeds. An experiment tests whether we can improve it. Production reveals whether that improvement holds under real constraints. These should not be separate activities.

  1. 01EvaluateMeasure capability
  2. 02OptimizeTest an improvement
  3. 03ServeRun real workloads
  4. 04MeasureCheck the result
Measured results feed the next evaluation. The loop continues.

CodeSOTA provides model evidence. Research explores better training and inference methods. Fabryka brings useful improvements into production.

The goal is not more projects. It is a system in which each piece makes the others more valuable.

02 / Evidence-based routing

Quality is specific to the task.

There is no single best model for every request. A model that writes excellent code may be the wrong choice for extracting fields from a Polish invoice. A large reasoning model may be unnecessary for classifying a support ticket.

A cheap answer is not cheap if someone has to correct it. We want model selection to follow evidence: task-level quality, cost, response time, and deployment constraints.

Use the smallest system that reliably completes the job. Escalate when the task demands more.

That is the direction for Fabryka Router: from selecting an available model to selecting a system supported by measured capability.

03 / Built around real work

Polish is a workload,
not a label.

Supporting Polish means more than generating fluent sentences. It means understanding documents, following instructions, working with local terminology, using tools, and knowing when the available evidence does not support an answer.

We want evaluations that reflect that work:

An aggregate score is a starting point. What matters is whether a model can complete your task.

04 / Research discipline

Research has to earn its compute.

New ideas should start as small, controlled experiments—not expensive commitments. Our research direction is a scaling ladder: test cheaply, compare against a baseline, and scale the ideas whose improvements survive.

A better result on a small model is a reason to investigate. It is not proof that the same method will work at a larger scale.

The same standard applies to checkpoint selection, distillation, quantization, and model merging. We care about the resulting capability, cost, and reliability—not whether a technique is fashionable.

Research line

Room to experiment.

Small models, new training recipes, and controlled comparisons. Failed experiments are useful results.

Production line

Evidence before release.

Useful capabilities, measured tradeoffs, and reliable serving. An experiment is not a production promise.

05 / Specialist models

Not every task needs
a general-purpose model.

Much of an AI workflow consists of narrow, repeatable work: classification, extraction, reranking, document cleanup, and routing. Small specialist models may handle these steps more efficiently, leaving larger models for the parts that need them.

We want to build and adapt models around useful jobs—not parameter counts.

The question is not “Can this model replace a general assistant?” It is “Can this system complete this workload reliably, at a lower total cost?”

06 / Accountable data

Better systems need
accountable data.

Distillation and fine-tuning are only as trustworthy as the data behind them. Our standard should be traceable sources, documented permissions, duplicate checks, privacy review, and evaluation data kept separate from training.

Synthetic data needs the same scrutiny. A teacher’s answer is a candidate—not ground truth.

Production feedback also needs clear boundaries. Operational measurements can guide optimization without turning customer conversations into training data. Any use of customer content for research should require explicit permission and defined handling rules.

Improving the system must not come at the expense of the people using it.

07 / Open evidence

Publish the evidence.
Including the failures.

We want technical claims to come with enough context to inspect them: the model, hardware, configuration, workload, and measurement method.

When an optimization improves throughput but harms quality, both results matter. When an experiment fails, that is worth documenting too.

A public research log should show what we tried, what changed, and what did not work. Trust should come from reproducible results—not adjectives.

08 / Where we are going

What exists.
What comes next.

Next layer

Connect the evidence to decisions.

Connect evaluations more tightly to routing and model selection. Make training experiments comparable. Establish clear data-review gates. Publish the evidence behind improvements.

Research direction

Specialists that work together.

Specialist models, repeatable distillation, and customer-specific adaptation—working together as a more efficient system for real workloads.

This is a direction, not a claim that every part is already built.

The long-term view

The API is where you start.
The factory is what we’re building.

More useful work from open models. Better decisions from evidence. More intelligence from the hardware we run.

That is the vision for Fabryka.