Feedback on Laguna-S-2.1 long-form generation behavior

#9
by ChuckK1138 - opened

Hello Poolside Team,

First, thank you for making the Laguna models available to the community. I have been evaluating Laguna-S-2.1 extensively in a local deployment, and I wanted to share an observation that may be useful to your engineering team.

Overall, I have been very impressed with the model. Its reasoning, organization, writing quality, and especially its handling of citations are excellent. In many cases, I prefer its style to other open-weight models I have tested.

During extended testing, however, I consistently encountered a specific behavior during long-form responses.

The model generally produces an excellent answer and reaches what appears to be a natural conclusion. Instead of terminating generation at that point, it often continues generating additional text. The continuation gradually shifts away from the original task and enters what I would describe as a "free association" mode. Rather than stopping, the model continues expanding on loosely related concepts until generation eventually ends because of the context or token limit.

The important point is that this behavior does not occur throughout the response. The primary answer is typically coherent, well-structured, and complete. The issue appears only after the response has already reached a logical conclusion.

I experimented with different system prompts, including substantially simplifying them, to determine whether prompt complexity was contributing to the behavior. While prompt changes affected the model's overall behavior in some respects, they did not eliminate this specific long-generation failure mode. The phenomenon remained reproducible.

For comparison, I ran many of the same prompts against other open-weight models in the same environment. While every model has its own strengths and weaknesses, this particular "post-conclusion continuation" behavior appeared to be specific to Laguna-S-2.1 in my testing.

My test environment is:

Apple Mac Studio (M3 Ultra)
512 GB unified memory
Apple MLX inference backend
Ollama v0.32.5
Open WebUI v0.11.0
Local inference (no cloud services involved)

I am sharing this because I believe Laguna has tremendous potential, and I wanted to contribute a careful observation rather than simply report that "the model rambles." My impression is that the model generally recognizes how to produce a complete answer but occasionally fails to recognize that it has already completed the task, resulting in continued semantic expansion instead of terminating generation.

If additional examples, prompts, logs, or reproduction steps would be useful, I would be happy to provide them.

Thank you again for releasing these models and for your continued work on open-weight AI.

Best regards,

Charles Kimble

cdkimble@hotmail.com
Secure AI Systems, LLC

Sign up or log in to comment