Our new website is still in development. Some pages and details are being finalised.

Open models.

Published Evaluations.

Stated Limits

Crane's research output is the reason technical evaluators take the deployment work seriously: models released openly, benchmarks other institutions run, and limitations written down rather than left out.

Research areas

Crane's research focuses on language, speech, cultural evaluation, clinical and educational adaptation, agricultural intelligence and efficient deployment. We publish selected models, datasets, benchmarks and technical notes so others can inspect, reproduce and build on the work.
Efficient deployment
Quantisation, runtime optimisation and device testing for constrained compute environments.
Benchmarks and evaluation
Cultural, linguistic, pedagogical and domain evaluations designed to reveal failures hidden by broad global benchmarks.
Agriculture AI
Weather-informed advisory, crop diagnosis, and local-language agricultural information systems for low-connectivity environments.
Education AI
Foundational literacy and numeracy models, pedagogical content knowledge benchmarks, and teacher-facing tools for local languages.
Health AI
Clinical adaptation, safety filtering and on-device deployment of medical model families.
Speech
Streaming ASR and text-to-speech systems for local-language interfaces.
Language models
Compact and adaptable model families for African languages.
Crane-Med

Uganda clinical context

Research | Deployment component
Luganda FastPitch TTS

Luganda

Research | Component
Ganda-Gemma LiteRT

Luganda

Released
BridgeFLN

Luganda, education

Released
Swahili-Gemma-1B

Swahili

Released research model
Ganda-Gemma-1B

Luganda

Released research model
MedASR

Medical speech workflow

Research | Component
Crane Nemotron 4B

Seven East African languages

Research checkpoint
Crane Nemo ASR v1.2

Luganda, Shona, Swahili; English retained

Research release
Luganda Medical Speech Corpus
Open clinical speech corpus
Licence TBC — open
Afri-Aya
13 languages
VERIFY award wording
Pedagogy Benchmark Multilingual
Teacher-training exam questions translated into African languages
VERIFY
ELL benchmark
100 Luganda linguistics items
VERIFY
PCK benchmark
100 Luganda pedagogy MCQs
VERIFY
Luganda FLN dataset
13,224 MCQs
Apache-2.0
WildJailbreak Africa
~299,283 samples
ODC-BY-1.0
UCCB
1,039 items
CC BY-NC-SA 4.0

Technical contributions

Custom Bantu tokenisers reducing the compute penalty from 4.51× to 1.9× · 17-model tokeniser evaluation using Bits Per Character · GGUF and llama.cpp quantisation for sub-6GB RAM devices · LLM-as-judge cultural and pedagogical evaluation methodology · multimodal pipelines combining on-device vision with reasoning.

Tokenisation

Cutting the compute penalty from 4.51× to 1.9×

Custom tokenisers built for Bantu languages, benchmarked against 17 models using bits per character to confirm the gain held across the field.

DEPLOYMENT & EVALUATION

Running well under 6GB, judged on more than accuracy

GGUF and llama.cpp quantisation for constrained devices, an LLM-as-judge for cultural and pedagogical quality, and multimodal pipelines pairing on-device vision with reasoning.

Model-card Template

Reusable. Every model published by Crane follows this structure.

Hero

[Model] · [version] · [status] · [release date]

Summary

What it does, for which languages and tasks, why it exists

Architecture

Base model, adaptation approach, size — no proprietary recipes

Languages & domains

Exactly what was evaluated; distinguish trained, supported, experimental

Benchmarks

Method, dataset, device/runtime, comparator, date

Intended use

Supported applications and expected human oversight

Limitations

Quality gaps, dialect and geography limits, safety boundaries, unsupported tasks

Deployment

Weights, runtime, device/cloud requirements, examples

Licence & attribution

Exact licence, base-model attribution, citation

Resources

Hugging Face · GitHub · technical note · contact

Peer-reviewed research

Our work is published, cited and inspectable.

Research Title

ACM CHI 2026

Presented in Barcelona · Extended Abstracts (CHI EA '26) - Selected from 198 workshop proposals; 70 accepted.

Authors:

Kato Steven Mubiru and Bakunga Bronson

(co-authors & co-organisers)

DOI: 10.1145/3772363.3778692

VERSIONING & ATTRIBUTION

How we keep our models honest and transparent

We believe in radical transparency. Every model version is tracked, every contributor is credited, and every deployment is documented. No silent updates, no hidden data.

The Transparency Manifesto

Watch how we build trust into every line of code.

Versioning

Every model update is tracked and immutable.

Attribution

Contributors are credited for every line of code.

Documentation

Every deployment is documented for auditability.

Hugging Face

GitHub