Over the past twelve months, large language models (LLMs) have emerged as the primary focus of debate in generative AI. Small language models (SLMs) have quietly generated attention among professionals amid the boom of proprietary LLMs like ChatGPT.
Microsoft announced the introduction of Phi-2 earlier this month. Microsoft Phi-2 model is a 2.7 billion parameter SLM with “outstanding” language understanding and reasoning abilities. With less than 13 billion parameters, this model has attained state-of-the-art performance and can outperform models 25 times larger.
Microsoft loves SLMs
In addition to lightweight model training techniques like Phi and Phi 1.5, Microsoft launched Ocra earlier this year. It is an open-source model that takes Vicuna as inspiration and can replicate and learn from GPT- 4-size LLMs.
Additionally, Nadella revealed “models-as-a-service” solutions at Ignite 2023. These offerings give businesses access to a variety of open-source models. This includes those from Mistral and LLaMA 2 on platforms like Hugging Face.
Also, Phi-2, which businesses can access through the Azure AI catalog, is now a competitor for the LLaMA model series. Microsoft stated earlier this year that Phi-1.5 beats the 7-billion parameter model of LlaMA 2 on several benchmarks.
Phi-2 and the SLM Revolution
Since they can function more economically and utilize less processing power than LLMs to develop insights, SLMs are becoming increasingly popular in the generative AI market.
Phi-2 trained on 96 Nvidia A100 GPUs in roughly 14 days, compared to GPT-4’s alleged 90–100 day training period on 25,000 A100 GPUs. It hasn’t quite matched GPT-4’s performance but outperformed bigger models on several benchmarks. One of the benchmarks includes the elimination of major Generative AI accessibility issues.
More precisely, it performs better than Mistral 7B and Llama-2 models in BBH, common sense thinking, coding, math, and language understanding (Llama 2 only). Furthermore, it has surpassed Gemini Nano 2 on several benchmarks, like MBPP, MMLU, BoolQ, and BBH.
Given that Phi-2 has matched or even surpassed Llama 2 70B on some benchmarks, it is clear that SLMs—even those with more parameters—are democratizing AI development.
Phi-2: Small AI-Language Models Benefits
Here are a few small AI language model benefits that have been revolutionizing the language processing and generative AI landscape:
- Computational Efficiency: Due to their reduced computational requirements for training and inference, small language models are a more practical choice for users with constrained resources or for devices with less powerful processors.
- Quick Inference: Smaller models have quicker inference times. It makes them ideal for real-time applications where low latency is critical to the outcome.
- Resource-Friendly: Small language models are optimized by design to use less memory. It makes them perfect for implementation on devices with limited resources, such as Edge or smartphone devices.
- Energy Efficient: Unlike risks of open-sourced AI models during training, small models are more energy-efficient during training and inference. That’s because of their smaller size and lower complexity. This makes them suitable for applications where energy efficiency is a key consideration.
- Reduced Training Time: Unlike their bigger counterparts, training smaller models takes less time, which is a big advantage when quick model iteration and deployment are crucial.
- Improved Interpretability: Smaller models are frequently easier to read and understand. This is especially important for applications where transparency and model interpretability are critical, like legal or medical domains.
- Economic Solutions: Smaller models are less costly to train and implement in terms of time and computational resources. They are a good option for people or organizations on a tight budget because of their accessibility.
- Designed for Particular Domains: A more compact and appropriate language model than a vast, all-purpose one might be used in some specialized or domain-specific applications.
Phi-2 Model and Training Data
Microsoft has indicated that the quality of the SLM’s training data is one of the main factors influencing its success in the Phi-2 example.
Microsoft Phi-2 model is trained with “text-book quality” training data. It combines synthetic datasets to educate the model on general knowledge, common sense reasoning, and more. Then, this artificial data combines with web data that has been “filtered based on educational value and content quality.”
Ending Note
The performance of the Microsoft Phi-2 model on reasoning tasks compared to Llama 2 70B indicates that although SLMs are still far from matching the capabilities of leading LLMs such as GPT-4, the gap is narrowing. SLMS is a viable option for businesses looking to use generative AI more cheaply and efficiently in terms of computing.
For more informative and valuable content on technology and generative AI, keep exploring the blogs at Storytelling With Charts.




















