H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

Most visual document retrievers in production today are hand-me-downs. ColPali and the models that followed it take a generative vision-language model and repurpose it as an encoder. The result still carries a separately pretrained vision tower and a causal decoder that never generates a token. That is parameter and compute overhead for a task that…

Read Full News

My Brief Summer Fling With Siri AI

When I first tested the developer beta of Apple’s Siri AI by taking it as my tour guide around San Francisco, I was convinced the overhauled smartphone assistant would be an “everything tool” on the iPhone. The answers felt reliable enough and more helpful than previous versions of Siri. Also, the chatbot-style app for Siri…

Read Full News

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

A team of researchers from UC Berkeley have released CUA-Lite, an open platform for computer-use agents (CUAs). The argument behind it is infrastructural rather than model-centric: training and benchmarking a CUA requires four pieces: agents, environments, traces, and a framework to evaluate and train them and all four are currently fragmented across separate repositories with…

Read Full News

Seattle Occasions and Newsday are the newest publications to sue OpenAI and Microsoft

Two extra information organizations are suing OpenAI and Microsoft over the supposed use of their journalism to coach AI. A lawsuit filed by The Seattle Occasions and Newsday argued that with the arrival of AI, the journalism trade may change into “damaged past restore.” The lawsuit described generative AI as “a snake consuming its personal…

Read Full News