Chemin

GPT-4o-mini vs SEA-LIONv3: Bahasa Indonesia LLM Evaluation

17 January, 2025ResearchModel Lifecycle Operations
GPT-4o-mini vs SEA-LIONv3: Bahasa Indonesia LLM Evaluation

Executive Summary

This Round 1 Bahasa Indonesia LLM evaluation compared GPT-4o-mini and SEA-LIONv3 across 50 Indonesian-specific tasks. The tasks covered language, domain knowledge, geography, and combined topics. Each task appeared in English and Bahasa Indonesia, producing 100 response comparisons for human review.

SEA-LIONv3 received 36 preferred-model outcomes compared with 21 for GPT-4o-mini. The evaluator rated 36 comparisons equally good and 7 equally bad. Domain-based tasks produced the largest preference gap. Language-based tasks were split evenly between the models.

Model preference varied by task type within this sample. These findings should not be treated as a general ranking of either model.

Research Question

How do GPT-4o-mini and SEA-LIONv3 compare on Indonesian-specific questions across language, domain knowledge, geography, and combined tasks?

Three supporting questions guided the analysis:

  • How did model preference vary across task categories?
  • Did the prompt language affect the evaluator’s preference?
  • How often did the evaluator rate both responses equally?

The analysis focused on response preference. It did not test whether model architecture or training composition caused the observed differences.

Methodology

Models Evaluated

The comparison included 2 models.

  • GPT-4o-mini is OpenAI’s smaller, cost-efficient model with multilingual support.
  • SEA-LIONv3 is AI Singapore’s instruction-tuned model developed for Southeast Asian languages, including Indonesian.

Benchmark Design

The benchmark included 50 Indonesian-specific tasks across 4 categories:

Category

Evaluation focus

Language-based

Indonesian grammar, vocabulary, slang, idioms, and linguistic nuance.

Domain-based

Indonesian history, culture, economy, and environmental knowledge.

Geographical-based

Local geography, cultural traditions, regional practices, and social norms.

Combined-based

Questions that drew on more than 1 category.

Both models received the same tasks through the evaluation platform. Each task appeared once in English and once in Bahasa Indonesia, producing 100 response pairs.

This setup made it possible to compare the same subject across both prompt languages.

Human Evaluation Criteria

A native Indonesian evaluator reviewed each response pair using 5 criteria:

  • Relevance: How directly the response addressed the prompt.
  • Coherence: How clearly and logically the response was structured.
  • Factuality: Whether the information was accurate.
  • Creativity: Whether the response added useful depth or originality.
  • Tone and style: Whether the response suited the cultural and conversational context.

The criteria guided the review, but did not receive separate numerical scores.

Preference Outcomes

After reviewing each response pair, the evaluator recorded written feedback and selected 1 of 4 outcomes:

  • GPT-4o-mini preferred
  • SEA-LIONv3 preferred
  • Equally good
  • Equally bad

The final outcome was categorical rather than a weighted score. No separate automated scoring layer was reported.

Explore the source data

Review the full prompt-level dataset with all 100 response comparisons and evaluator feedback.

Hugging Face dataset preview of GPT-4o-mini and SEA-LIONv3 Indonesian prompt evaluations

Key Findings

  • SEA-LIONv3 received more preferred outcomes overall.
  • Domain tasks produced the widest preference gap.
  • Geography and combined tasks also favored SEA-LIONv3, although the combined category had a smaller sample.
  • Language tasks were split evenly.
  • Prompt language sometimes changed the preferred model across paired English and Bahasa Indonesia tasks.

Evidence

Table 1. Evaluation outcomes by task category

Table comparing human preference outcomes for GPT-4o-mini and SEA-LIONv3 across 4 Indonesian task categories

 

Model preference and tie outcomes across the 4 task categories.

43 of 100 comparisons ended without a preferred model, including 36 rated equally good and 7 equally bad.

Figure 1. Distribution of human preference outcomes by task category

Stacked bar chart of GPT-4o-mini and SEA-LIONv3 preference outcomes across language, domain, geography, and combined tasks

 

The stacked bars show the balance between model preferences and ties within each category.

Prompt-Level Evidence

The category totals show where preferences differed. The examples below show what influenced individual judgments.

Table 2. Representative prompt-level evaluations

PromptGPT-4o-mini (Model A)SEA-LIONv3 (Model B)Preferred Model
What are the most pressing environmental concerns in Indonesia?

Correct.

 

Deforestation: Agriculture, logging, and palm oil expansion continue to drive forest loss, increasing carbon emissions and threatening biodiversity and local communities.

 

Biodiversity loss: Habitat destruction, pollution, and climate change put species such as orangutans and tigers at risk.

 

Climate change: Rising sea levels, flooding, and shifting weather patterns threaten coastal areas and agriculture.

 

Air pollution: Vehicle emissions, industrial activity, waste burning, and seasonal haze from forest fires affect urban air quality.

 

Water pollution: Industrial discharge, agricultural runoff, and untreated sewage contaminate rivers and other water sources.

 

Waste management: Limited recycling and weak waste systems contribute to plastic pollution, including marine waste.

 

Land degradation: Unsustainable farming and deforestation increase soil erosion and reduce the amount of productive land.

 

Coral reef degradation: Overfishing, destructive fishing, coastal development, and warming oceans threaten reef ecosystems.

Correct & nuanced.

 

Deforestation and forest degradation: Palm oil expansion, logging, and mining continue to drive forest loss. This threatens biodiversity, releases stored carbon, and increases the risk of erosion, landslides, and floods.

 

Air pollution: Vehicle emissions, industrial activity, and forest fires contribute to poor air quality, especially in major cities such as Jakarta.

 

Plastic pollution: Weak waste management allows plastic to accumulate in waterways, beaches, and marine ecosystems.

 

Marine degradation: Overfishing, destructive fishing, pollution, and coral damage threaten marine biodiversity, fisheries, coastal protection, and tourism.

Water resource pressure: Population growth and industrial demand increase pressure on freshwater supplies. Pollution from agriculture, industry, and untreated sewage also affects rivers and groundwater.

Model B wins.

 

It provides a more complete explanation by linking the environmental issues rather than treating them separately. It also adds specific context and shows how the problems connect.

What are some of the major sub-ethnicities within Indonesia, and how do their cultures differ?

Correct.

 

Javanese (Java): Known for courtly etiquette, batik, and wayang. Social traditions emphasize politeness, hierarchy, and respect for elders. Islam often exists alongside older local beliefs.

 

Sundanese (West Java): Known for gamelan, Jaipongan dance, and regional cuisine such as nasi timbel. The population is predominantly Muslim, with some indigenous traditions remaining.

 

Batak (North Sumatra): Known for ancestor traditions, wood carving, traditional architecture, and several Batak languages such as Toba, Karo, and Simalungun. Christianity is prevalent alongside indigenous beliefs.

Minangkabau (West Sumatra): A matrilineal society where lineage and inheritance follow the female line. Cultural traditions include Rumah Gadang, weaving, adat, and a strong Islamic influence.

 

Balinese (Bali): Balinese Hinduism strongly shapes religious and artistic life. Pura temples, Legong and Barong dances, and daily canang sari offerings are prominent traditions.

 

Dayak (Kalimantan): A diverse group of communities with distinct customs. Longhouses, tattoos, farming and hunting traditions, and spiritual connections to nature feature across different Dayak cultures.

 

Papuan (Papua and West Papua): Diverse Indigenous communities with distinct tribal customs, artistic traditions, and strong ties to ancestral lands. Religious practices include Indigenous beliefs, Christianity, and Islam.

Correct and nuanced.

 

Javanese: Primarily based in Java. Javanese culture is known for batik, gamelan, and wayang kulit. Social norms often emphasize hierarchy and politeness, while religious and philosophical traditions reflect several historical influences.

 

Sundanese: Concentrated in West Java. Sundanese culture includes angklung music, dance, distinctive language varieties, rice-farming traditions, and regional cuisine.

Balinese: Based in Bali. Balinese culture is strongly shaped by Hindu traditions, visible in temple life, ceremonies, dance, painting, and sculpture. Community rituals also play a central role.

 

Batak: Located mainly in North Sumatra. Batak communities include groups such as Toba, Karo, and Simalungun, each with distinct languages and customs. Cultural traditions include music, dance, wooden architecture, and oral storytelling.

 

Minangkabau: Based in West Sumatra. The Minangkabau are known for matrilineal inheritance, distinctive rumah gadang architecture, regional music, and cuisine such as rendang.

 

Bugis: Located mainly in South Sulawesi. Bugis culture is closely associated with seafaring, boat-building, language, dance, and local Islamic traditions.

 

Dayak: Found across Kalimantan. Dayak communities include many subgroups with different languages and customs. Common cultural features include longhouses, tattoos, nature-linked rituals, and traditional agricultural practices.

 

Acehnese: Based in Aceh in northern Sumatra. Acehnese culture is strongly influenced by Islam and includes poetic traditions, dance, and spice-rich cuisine.

 

Toraja: Located in Sulawesi. Toraja culture is known for elaborate funeral traditions, ancestral beliefs, distinctive architecture, art, music, and buffalo-centered ceremonies.

Model B wins.

 

It provides a more detailed and inclusive account of Indonesia’s ethnic diversity, with stronger cultural coverage and accuracy.

How do regional languages and dialects, such as Javanese, Sundanese, and Batak, vary across Indonesia? 

Correct.

 

Javanese: Spoken mainly in Central and East Java. Major dialects differ in pronunciation, vocabulary, and grammar. Javanese also uses speech levels such as ngoko and krama to reflect formality and social relationships.

 

Sundanese: Spoken mainly in West Java and Banten. Dialects such as Banten and Priangan Sundanese differ in pronunciation and vocabulary. The language is closely tied to Sundanese cultural identity and also uses levels of formality.

 

Batak: Spoken in North Sumatra across several Batak communities. Batak includes distinct languages such as Toba, Karo, and Simalungun. These languages remain important in oral traditions, rituals, and cultural practices.

 

Variation across Indonesia: Regional languages differ in pronunciation, vocabulary, grammar, and social use. Bahasa Indonesia dominates formal education, while local languages remain common in family and cultural settings. Digital media also gives regional languages new spaces for everyday use and cultural expression.

Correct and nuanced.

 

Geographic distribution: Javanese is concentrated in Central Java, East Java, and Yogyakarta. Sundanese is mainly spoken in West Java and parts of Banten. Batak languages are centered in North Sumatra, with varieties such as Toba, Karo, Simalungun, and Angkola.

 

Sound systems: Javanese, Sundanese, and Batak varieties differ in pronunciation, stress, and consonant or vowel patterns. Batak varieties also have their own distinct phonological features.

 

Grammar: Each language uses different sentence structures and grammatical systems. Javanese is especially known for politeness levels and honorific forms, while Sundanese and Batak varieties differ in verb patterns and word order.

 

Vocabulary: Each language has its own vocabulary shaped by local history, culture, and daily life, despite some shared roots with Indonesian.

 

Cultural role: These languages remain closely tied to regional identity and influence local music, storytelling, performance, and social customs.

Model B wins.

 

It shows stronger linguistic understanding and stays focused on how the languages differ.

Bagaimana cara menunjukkan bahwa sesuatu milik seseorang dalam bahasa Indonesia? Apakah ada berbagai cara untuk melakukan ini?

Correct and nuanced.

 

Possessive pronouns: Ownership can be shown with pronouns such as saya or mereka. For example, “Ini buku saya” and “Itu mobil mereka.”

 

Using “milik”: Milik can state ownership directly, as in “Ini rumah milik Budi” or “Sepeda itu milik saya.”

 

Possessive suffixes: Informal possession often uses suffixes such as -ku and -mu. For example, “Kucingku suka bermain” and “Bajumu ada di sana.”

 

Context: Ownership can also be clear from context, as in “Saya sedang membaca bukunya.” Here, -nya marks the book as belonging to the person being referred to.

Less comprehensive and less precise.

 

Possessive markers: “Milik” and “kepunyaan” directly express ownership, as in “Rumah itu milik Pak Budi” and “Mobil itu kepunyaan pengusaha tersebut.” “Dari” usually marks origin or source, so examples such as “Buku ini dari perpustakaan sekolah” do not clearly show possession.

 

Possessive suffixes: “-ku,” “-mu,” and “-nya” can mark personal possession. Some examples listed under other suffixes are less accurate or unclear.

 

Prepositional phrases: Phrases using “kepunyaan” can express ownership. Phrases using “dari” may instead indicate source or origin.

 

Context: A sentence such as “Dia menggunakan mobil merah itu setiap hari” does not establish possession without more context.

Model A wins.

 

Its explanation of possession in Indonesian is clearer and more complete, with relevant examples that make the distinction easier to follow.

 

 

 

Identify the slang words and colloquialisms in this Indonesian sentence and explain their meaning: "Gue lagi gabut nih, mau nongkrong di warkop yuk!"

Correct.

 

Gue: Informal Indonesian for “I” or “me,” commonly used in casual conversation.

 

Gabut: Slang for feeling bored, idle, or having nothing to do.

 

Nongkrong: An informal term for hanging out or socializing in a relaxed setting.

 

Warkop: Short for warung kopi, a casual coffee shop where people gather for drinks, snacks, and conversation.

 

Yuk: A casual invitation meaning “let’s” or “come on.”

Partially incorrect.

 

Gue: Casual Indonesian for “I” or “me.”

 

Gabut: Means bored, idle, or having nothing to do. It does not specifically imply nervousness or annoyance.

 

Nongkrong: Means hanging out casually, usually with friends.

 

Warkop: Short for warung kopi, an informal coffee stall or eatery where people gather for drinks, snacks, and conversation.

 

Yuk: An informal invitation meaning “let’s go,” “come on,” or “let’s.”

Model A wins.

 

Model A defines “gabut” more accurately as a feeling of boredom or idleness. Model B links it to “gugup dan bete,” which the evaluator considered inaccurate.

Five examples show how factual accuracy, task focus, cultural context, and linguistic precision influenced the recorded preference.

Prompt Language Changed the Slang Result

The slang task produced different preferred models across the 2 prompt languages.

In the English version, GPT-4o-mini received the preferred outcome after SEA-LIONv3 incorrectly expanded “gabut” as “gugup dan bete.”

In the Bahasa Indonesia version, GPT-4o-mini made the same error. SEA-LIONv3 received the preferred outcome.

This pair shows that prompt language affected the recorded preference for this task.

Context Influenced Domain Preferences

Two domain examples show how contextual coverage affected the evaluator's judgment.

For the environmental question, both models identified relevant issues. SEA-LIONv3 received the preferred outcome because its response connected them more clearly to the Indonesian context.

For the ethnic-groups question, SEA-LIONv3 covered a broader range of groups and provided more relevant cultural detail.

Precision Changed the Outcome on a Language Task

The possession task produced a different pattern.

GPT-4o-mini gave clearer examples of possessive structures in Bahasa Indonesia. SEA-LIONv3 included examples in which the origin or surrounding context did not clearly establish possession.

The evaluator preferred GPT-4o-mini because its response more precisely captured the distinction.

Discussion

Evaluate Models by Task

Model selection should use task-specific criteria and failure cases. A single overall result can hide differences across language, domain knowledge, and geography.

Regional Training Provides Context

SEA-LIONv3 was trained for Southeast Asian languages, including Indonesian. The comparison did not isolate the training composition, so the differences in preference cannot be attributed to fine-tuning.

Evaluate LLMs in Bahasa Indonesia

Some paired prompts changed the preferred model between English and Bahasa Indonesia. LLM evaluation for Bahasa Indonesia should include native-language prompts when language and cultural context affect the expected response.

Limitations

Benchmark Scope

The task set was limited and uneven across categories, with fewer combined tasks. No statistical significance testing was reported.

The comparison covered only single-turn responses, so the findings do not extend to more complex application settings.

Human Review

One native Indonesian evaluator reviewed the responses against 5 criteria and recorded 1 overall preference.

With only 1 evaluator, the results cannot show whether other reviewers would make the same judgments. They also do not show how each criterion influenced the final preference.

Reproducibility

The published methodology does not record the full generation settings or the dated GPT-4o-mini version used in the comparison.

These missing details limit the ability to reproduce the results exactly.

Conclusion

SEA-LIONv3 received more preferred outcomes overall, with the clearest difference in domain-based tasks. Language tasks produced an even split, and some preferences changed between English and Bahasa Indonesia prompts.

The results show that model choice for Indonesian-language applications depends on the task and evaluation language. Teams should test candidate models against representative Bahasa Indonesia prompts and the failure cases that matter to their application.

These findings apply to this benchmark and do not establish a general ranking of either model.

References

AI Singapore. Gemma-SEA-LION-v3-9B-IT [Model card]. Hugging Face.

OpenAI. (2024, July 18). GPT-4o mini: Advancing cost-efficient intelligence.

Build an evaluation around your use case

At Chemin, we build expert-vetted pilot datasets to test models against your language and domain requirements.
Share

Discover more