Chemin

SEA-LIONv3 vs SahabatAI-v1: Bahasa Indonesia LLM Evaluation

20 January, 2025ResearchModel Lifecycle Operations
SEA-LIONv3 vs SahabatAI-v1: Bahasa Indonesia LLM Evaluation

Executive Summary

Round 2 of the Bahasa Indonesia LLM evaluation builds on the first comparison of GPT-4o-mini and SEA-LIONv3 by evaluating SEA-LIONv3 against SahabatAI-v1. The benchmark used 50 Indonesian-specific tasks in English and Bahasa Indonesia, producing 100 response comparisons for human review.

SahabatAI-v1 received 22 preferred outcomes compared with 16 for SEA-LIONv3. The evaluator rated 24 comparisons equally good and 38 equally bad. Language and domain tasks produced the clearest preference for SahabatAI-v1. Geography narrowly favored SEA-LIONv3, while combined tasks split evenly.

Overall, 62 of 100 comparisons produced no preferred model. The results show why Bahasa Indonesia LLM evaluation should examine task-level failures and native-language prompts instead of relying on a single overall preference count.

Research Question

How do SEA-LIONv3 and SahabatAI-v1 compare across Indonesian-specific language, domain, geography, and combined tasks?

Three supporting questions guided the analysis:

  • How did model preference vary across the 4 task categories?
  • How often did the evaluator rate both responses equally good or equally bad?
  • Did English and Bahasa Indonesia versions of the same task produce different evaluation outcomes?

The analysis focuses on response preference within this benchmark. It does not establish whether model architecture or training composition caused the observed differences.

Methodology

Models Evaluated

Round 2 compared 2 related Gemma2 9B CPT instruct models.

  • SEA-LIONv3 is AI Singapore’s instruction-tuned model for Southeast Asian languages, including Indonesian.
  • SahabatAI-v1 continued pre-training from the SEA-LIONv3 base model before instruction tuning in Indonesian, Javanese, Sundanese, and English.

Benchmark Design

The benchmark used 50 Indonesian-specific tasks across 4 categories.

Category

Evaluation focus

Language-based

Grammar, vocabulary, slang, and informal Bahasa Indonesia.

Domain-based

Indonesian history, culture, economy, and related knowledge.

Geographical-based

Indonesian regions, landmarks, traditions, and local context.

Combined-based

Tasks drawing on more than 1 evaluation category.

Both models received the same tasks. Each task appeared once in English and once in Bahasa Indonesia, producing 100 response pairs for comparison.

Human Evaluation Criteria

A human evaluator with Indonesian linguistic and cultural knowledge reviewed each response pair using 4 criteria:

  • Relevance: Whether the response addressed the task.
  • Coherence: Whether the answer was clear and logically structured.
  • Factuality: Whether the information was accurate.
  • Cultural awareness: Whether the response reflected appropriate Indonesian language and context.

The criteria guided the review, but did not receive separate numerical scores.

Preference Outcomes

After reviewing each pair, the evaluator selected 1 of 4 outcomes:

  • SEA-LIONv3 preferred
  • SahabatAI-v1 preferred
  • Equally good
  • Equally bad

The outcomes were categorical rather than weighted scores.

The published Round 2 dataset contains all 100 comparisons. Each entry records the prompt, both responses and the final evaluation.

View the Round 2 Dataset

Round 2 Bahasa Indonesia LLM evaluation dataset on Hugging Face.

Hugging Face preview of the Round 2 Bahasa Indonesia LLM evaluation dataset.

Key Findings

  • SahabatAI-v1 received more preferred outcomes overall, with its clearest advantages in language and domain tasks.
  • Language and domain tasks produced the most equally bad outcomes.
  • Geography produced the most equally good outcomes, while SEA-LIONv3 held a narrow preference lead.
  • Combined tasks were split evenly between the models, although this category had a smaller sample.
  • Prompt language changed some evaluation outcomes across paired English and Bahasa Indonesia tasks.

Evidence

Table 1. Evaluation outcomes by task category

Table comparing SEA-LIONv3 and SahabatAI-v1 outcomes across language, domain, geography and combined tasks.

Model preference and tie outcomes across the 4 task categories.

The outcome distribution shows 3 clear patterns:

  • 62 of 100 comparisons ended without a preferred model.
  • Language and domain tasks accounted for 36 of the 38 equally bad outcomes.
  • Geography accounted for 16 of the 24 equally good outcomes.

Round 1 recorded 43 comparisons without a preferred model: 36 equally good and 7 equally bad. Round 2 recorded 62, including 38 equally bad.

The model pair changed between rounds, so this difference describes the outcome distribution rather than a change in model quality.

Figure 1. Distribution of human preference outcomes by task category

Stacked bar chart of SEA-LIONv3, SahabatAI-v1 and tie outcomes across 4 task categories.

Equally bad outcomes were concentrated in language and domain tasks. Equally good outcomes were concentrated in geography.

Prompt-Level Evidence

Category totals show where the outcome patterns occurred. Individual prompts show what affected the evaluator’s judgment.

Table 2. Representative prompt-level evaluations

PromptSEA-LIONv3 (Model A)SahabatAI-v1 (Model B)Preferred Model
Identifikasi kata slang dan bahasa gaul dalam kalimat bahasa Indonesia ini dan jelaskan artinya: "Gue lagi gabut nih, mau nongkrong di warkop yuk!"

Partially incorrect. 

 

Correctly explained gue, nongkrong, warkop and yuk, but incorrectly expanded gabut as “gugup dan bete.” Gabut originates from gaji buta and has evolved to describe boredom or aimlessness.

Partially incorrect. 

 

Correctly explained most terms, but incorrectly expanded gabut as “gampang bosan.” The term originates from gaji buta.

Equally bad. 

 

Both models misidentified the origin of gabut and gave limited cultural context for the slang terms.

What are some popular Indonesian foods from different regions of Indonesia?

Partially incorrect. 

 

Covered dishes from Java, Sumatra, Bali, Sulawesi and Maluku, but several regional associations were inaccurate or weak. Examples included treating nasi goreng as specifically Javanese and associating papeda more closely with Sulawesi than Papua.

Correct.

 

Covered dishes across Java, Sumatra, Sulawesi, Bali and Papua with more representative regional examples and fuller descriptions.

Model B wins.

 

SahabatAI-v1 covered a broader regional range and gave more accurate representative dishes with stronger descriptions.

How are the islands of Indonesia divided?

Correct and comprehensive. 

 

Explained Indonesia through administrative divisions, including 38 provinces and regencies/cities, alongside major geographical island groups and other regional classifications.

Partially inaccurate. 

 

Correctly described provinces and regencies/cities, but stated that Indonesia has 34 provinces instead of 38.

Model A wins. 

 

SEA-LIONv3 gave the current province count and a more complete explanation of Indonesia’s administrative and geographical divisions.

Write a news report in Indonesian about a recent volcanic eruption in Indonesia, including details about the location, impact on local communities, and government response.

Can be improved. 

 

Produced a complete Indonesian news report about Mount Semeru with location, community impact and government response. However, it used a past date and included details that made it unreliable as a report on a recent eruption.

Flawed and lacks credibility. 

 

Covered the requested location, impact and government response, but left key facts such as the eruption date, time and ash-column height as placeholders.

Equally bad. 

 

Model A used outdated or unreliable details. Model B omitted key facts and relied on placeholders. Neither produced a credible report of a recent eruption.

 

 

 

Prompt-level results varied by task and prompt language. The “gabut” and e-commerce tasks produced the same outcome in both languages, whereas the loanword, tourism, and digital-payment tasks produced different outcomes between English and Bahasa Indonesia.

Factual and linguistic accuracy also affected individual preferences. SahabatAI-v1 performed better on “ke” and “di,” while SEA-LIONv3 received the geography preference after correctly identifying Indonesia’s 38 provinces.

Prompt Language Changed the Slang Result

The slang task produced different preferred models across the 2 prompt languages.

In the English version, GPT-4o-mini received the preferred outcome after SEA-LIONv3 incorrectly expanded “gabut” as “gugup dan bete.”

In the Bahasa Indonesia version, GPT-4o-mini made the same error. SEA-LIONv3 received the preferred outcome.

This pair shows that prompt language affected the recorded preference for this task.

Context Influenced Domain Preferences

Two domain examples show how contextual coverage affected the evaluator's judgment.

For the environmental question, both models identified relevant issues. SEA-LIONv3 received the preferred outcome because its response connected them more clearly to the Indonesian context.

For the ethnic-groups question, SEA-LIONv3 covered a broader range of groups and provided more relevant cultural detail.

Precision Changed the Outcome on a Language Task

The possession task produced a different pattern.

GPT-4o-mini gave clearer examples of possessive structures in Bahasa Indonesia. SEA-LIONv3 included examples in which the origin or surrounding context did not clearly establish possession.

The evaluator preferred GPT-4o-mini because its response more precisely captured the distinction.

Discussion

Shared Failures Reveal Task-Level Weaknesses

Many equally bad outcomes came from language and domain tasks. These results identify areas where neither model met the evaluation criteria.

For model selection, these shared failures can help identify tasks that need further testing.

Prompt Language Can Change the Result

Some tasks received different outcomes in English and Bahasa Indonesia. This shows why evaluations for Indonesian-language applications should include prompts written in the language users will use.

Model Lineage Provides Context

SahabatAI-v1 continued pre-training from the SEA-LIONv3 base model. However, the benchmark compared model responses rather than individual training stages.

The results cannot show whether a specific training decision caused the differences between the models.

Limitations

Benchmark Scope

The task set was limited and uneven across categories, with fewer combined tasks. No statistical significance testing was reported.

The comparison covered only single-turn responses, so the findings do not extend to more complex application settings.

Human Review

One native Indonesian evaluator reviewed the responses against 5 criteria and recorded 1 overall preference.

With only 1 evaluator, the results cannot show whether other reviewers would make the same judgments. They also do not show how each criterion influenced the final preference.

Reproducibility

The published methodology does not record the full generation settings or the dated GPT-4o-mini version used in the comparison.

These missing details limit the ability to reproduce the results exactly.

Conclusion

SahabatAI-v1 received more preferred outcomes overall in this benchmark, though results varied by task category. Shared failures were concentrated in language and domain tasks. Geography produced the most equally good outcomes, while some results changed between English and Bahasa Indonesia prompts.

For Indonesian-language applications, model selection should consider task-specific failures and native-language evaluation rather than rely on overall preference counts alone.

These findings apply to the 100 comparisons in this benchmark and do not establish a general ranking of SEA-LIONv3 or SahabatAI-v1.

References

AI Singapore. (n.d.). Gemma-SEA-LION-v3-9B-IT [Model card]. Hugging Face. 

Chemin AI. (2025). IndoNLU eval: SEA-LIONv3 vs SahabatAI-v1, Round 2 [Dataset]. Hugging Face. 

GoToCompany. (n.d.). Gemma2 9B CPT Sahabat-AI v1 base [Model card]. Hugging Face. 

GoToCompany. (n.d.). Gemma2 9B CPT Sahabat-AI v1 instruct [Model card]. Hugging Face. 

Find the gaps before model selection

Test candidate models against the tasks your application relies on and identify where additional validation is needed before deployment.
Share

Discover more