Skip to content

FABRYKA AI / RESEARCH

Questions worth
measuring.

Three connected programmes, from learning with limited compute to measuring the systems we deploy. Current artifacts and planned studies are identified separately.

FAB-SLM-001 / Small Language Models

Scaling data mixtures and distillation for compact language models

Research programme · public model artifacts available

Research question
How do data composition, tokenization and teacher supervision affect model quality under a fixed parameter and training-compute budget?
Study design
Compare matched training runs with fixed evaluation splits; vary one factor at a time; record tokenizer, data revisions, tokens, seeds and compute. Report per-task results and limitations alongside aggregate metrics.
Scope & compute
Start with sub-150M models. Studies at 1B and 10B parameters are roadmap directions, subject to compute and data readiness.
Outputs & next steps
Published model repositories provide starting artifacts; the formal scaling study and its technical report are planned.

FAB-SYS-001 / Efficient Inference

Characterizing LLM inference on commodity GPUs

Research programme · formal comparative study planned

Research question
How do memory bandwidth, attention implementations, quantization and decoding strategies affect latency, throughput, memory use and output quality?
Study design
Compare the same model and workload across documented hardware and software configurations. Measure prefill, time to first token, decoding speed and latency distributions under specified concurrency. Evaluate quality separately from serving speed.
Scope & compute
Quantization, speculative decoding including DFlash, routing and commodity GPUs. Each result must state context length, batch size and load conditions.
Outputs & next steps
Production serving and token-probability tools support experiments. A controlled hardware comparison and technical report are planned.

FAB-DATA-001 / Polish AI Evaluation & Data

Polish data and reproducible AI evaluation

Research programme · public dataset and evaluation infrastructure

Research question
How do dataset composition and evaluation design affect our estimates of model capability in Polish, including document and OCR tasks?
Study design
Document data provenance, licensing, splits and scoring. Check contamination and reference quality. Use task-appropriate metrics and matched evaluation protocols; distinguish our measurements from externally reported results.
Scope & compute
Polish DynaWord, CodeSOTA benchmarks, Polish LLM evaluations and OCR evaluation methodologies.
Outputs & next steps
Polish DynaWord and CodeSOTA are public. Further evaluation studies and methodology reports form the research agenda.

PUBLICATION STANDARD

From hypothesis
to evidence.

Our intended reporting format connects a research question, hypothesis, baseline, method, experiments, results and limitations. Reports should identify authors, dates, funding and compute sources, and link the artifacts needed to inspect or reproduce the work.

Programme identifiers organize this research agenda. Completed experiments and published reports are listed only when a corresponding artifact is available.

Browse current outputs →