Burmese-Coder-4B: Fine-Tuning a Small Language Model for Burmese Coding

· ·

Burmese coding requires more than functional correctness

Low-resource code generation remains underexplored because most benchmarks focus on English. Burmese coding creates a dual challenge: a model must generate correct programs from Burmese instructions while keeping the surrounding natural-language response linguistically reliable.

In practice, a model can pass unit tests and still be unreliable for a native Burmese user. It may use unstable terminology, mix scripts, or drift into non-target languages around otherwise correct code. Functional correctness therefore does not fully measure the quality of a Burmese coding assistant.

Burmese-Coder-4B studies this gap through a Burmese code-generation model and a language-aware evaluation framework.

Featured on HackerNoon: Burmese-Coder-4B: A Burmese Coding LLM for Low-Resource Language AI covers the project for a broader developer audience.

The model and release

Burmese-Coder-4B is a coding-focused model adapted from the Gemma 3 4B family. It is designed for Burmese programming prompts and code generation, with evaluation that captures both program correctness and the quality of Burmese-language interaction.

The HackerNoon feature introduces the work to a broader audience. Hugging Face is the practical release location for developers who want to download, run, or evaluate the model.

Model adaptation and evaluation

The training process has two stages.

Stage 1: supervised fine-tuning

The first stage uses Burmese MBPP, a dataset of 974 programming tasks. Each example contains a Burmese instruction, a Python solution, and a Burmese explanation of the programming logic.

This stage teaches the model to respond to Burmese programming instructions with code and supporting Burmese-language content.

Stage 2: preference alignment

The second stage uses Direct Preference Optimization (DPO). Its goal is to reduce multilingual drift: responses that become unnecessarily mixed between Burmese and English, or that lose clarity while explaining technical concepts.

The model is not evaluated only on output code. The alignment stage targets more reliable Burmese-language responses around code, including reduced script mixing and more stable technical terminology.

For training settings, data preparation, and the complete experimental procedure, see the technical whitepaper.

How we evaluated it

A coding model should not be evaluated with a single metric. Burmese-Coder-4B uses a two-track evaluation design.

TrackQuestion
Functional correctnessDoes the generated code pass the unit tests?
Language-aware qualityIs the Burmese response fluent, instruction-faithful, semantically correct, terminologically stable, and free of mixed-script drift?

Functional correctness is measured with Pass@1 on Burmese HumanEval. Language-aware quality uses rubric-based judging across fluency, instruction following, semantic correctness, terminology quality, and mixed-language contamination.

Results

The central result is that Burmese-Coder-4B retained the base model’s functional score while improving language-aware quality for Burmese coding tasks.

ModelPass@1 (%)Rubric: DeepSeekRubric: Gemini
Burmese-Coder-4B62.03.4563.779
Gemma 3 4B base62.02.9393.203
Qwen 2.5 3B45.01.2202.526

The language-stability results are especially important:

  • Gemini mixed-language penalty: 0.69 -> 0.02
  • DeepSeek mixed-language penalty: 0.72 -> 0.09

Alignment improved linguistic reliability without reducing measured code correctness against the base model.

Example: a Burmese prompt to working Python

Burmese-Coder-4B demo output showing a Burmese coding prompt and generated Python response

Burmese-Coder-4B demo output: a Burmese prompt and the model’s generated Python response.

Limitations

Burmese-Coder-4B is an experimental research model. Generated code should be executed, tested, and reviewed before use in production. The current evaluation focuses on Python programming tasks and a limited benchmark set; broader language, framework, and human-evaluation studies remain important future work.

Conclusion

Burmese-Coder-4B is a Burmese code-generation project with language-aware evaluation. Its contribution is not only a fine-tuned model, but also a way to measure whether a low-resource coding assistant remains useful after it has generated the code.

The results suggest that targeted fine-tuning and preference alignment can improve linguistic reliability without sacrificing code-generation performance. For low-resource coding assistants, that means Pass@1 should be treated as necessary but not sufficient evidence of quality.

← All Projects