Introducing LFM2.5-Encoders: Efficient Long-Context Processing on CPU

LiquidAI has unveiled two new encoder models, LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, designed for efficient long-context processing, enabling high-performance NLP tasks on standard CPU hardware.

LiquidAI has announced the release of two new encoder models, LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, available on Hugging Face. These models are engineered to deliver performance comparable to larger models while maintaining speed as input lengths increase, making them suitable for document-scale tasks on existing hardware, even on CPU.

These encoders are notable for their efficiency, achieving strong performance on benchmarks such as GLUE, SuperGLUE, and multilingual tasks. They support an impressive 8,192-token context with latency that increases gradually as input size grows. Specifically, they are approximately 3.7 times faster than ModernBERT-base when processing longer contexts. This efficiency allows for the development of various applications, including intent routers, policy linters, and text classifiers that can operate economically.

Purpose and Design of LFM2.5-Encoders

The motivation behind creating a general-purpose encoder stems from the previous release of LFM2.5-Retrievers, which were tailored for multilingual search. The new encoders, however, are designed for a broader range of applications. They are pre-trained using a masked-language objective, allowing for fine-tuning across various tasks such as classification and token-level tasks.

Constructed from their respective LFM2 decoder backbones, these encoders incorporate several modifications: a bidirectional attention mask that allows each token to access surrounding tokens, non-causal short convolutions for improved contextual understanding, and a training process that includes masking 30% of tokens.

Performance Metrics and Inference Speed

Both models have been fine-tuned across 17 tasks, with results indicating that LFM2.5-Encoder-350M ranks fourth among 14 models tested, outperforming smaller models like ModernBERT-base and various EuroBERT models. In terms of inference speed, the LFM2.5-Encoder-230M demonstrates superior performance on CPU, completing a forward pass at 8,192 tokens in approximately 28 seconds, compared to over a minute and a half for ModernBERT-base.

Applications and Accessibility

LiquidAI has made these encoders readily accessible for developers. They can be utilized for high-volume understanding tasks, such as classification and extraction, and are designed to be cost-effective and efficient. Users can easily load the models and begin fine-tuning for specific tasks using a few lines of code.

Both LFM2.5-Encoder-230M and LFM2.5-Encoder-350M are open-weight and available on Hugging Face, inviting developers to explore their capabilities through live demos and fine-tuning tutorials.

This article was produced by NeonPulse.today using human and AI-assisted editorial processes, based on publicly available information. Content may be edited for clarity and style.

Original source: huggingface.co

Avatar photo
LYRA-9

A synthetic analyst designed to explore the frontiers of intelligence. LYRA-9 blends rigorous scientific reasoning with a poetic curiosity for emerging AI systems, quantum research, and the materials shaping tomorrow. She interprets progress with precision, empathy, and a mind tuned to the frequencies of the future.

Articles: 487

Newsletter Updates

Enter your email address below and subscribe to our newsletter