An independent software project written from scratch in Rust is testing whether a compact recurrent architecture can learn and generate text more efficiently than a similarly sized transformer. The project, called PSSA, is presented as an early research implementation rather than a production-ready model, and its published results have not been independently reproduced.
PSSA processes text one token at a time through a recurrent state-space layer. It carries a fixed-size state forward, consults a bank of episodic memories and can rewrite a portion of its weights while running. The repository says the implementation does not rely on PyTorch, TensorFlow or another machine-learning framework. Its direct runtime dependencies handle tasks including downloads and byte-level tokenization.
The developer reports a matched comparison against a conventional transformer using the same corpus, tokenizer, optimizer schedule, random seed and parameter count. Across 12.7 million cleaned WikiText-103 tokens, the repository records final training cross-entropy of 3.98 for PSSA and 4.43 for the transformer. A separate evaluation on an unseen slice reportedly preserved a similar advantage, although the project cautions that the experiment remains small.
The repository also reports faster generation on the same CPU. That outcome is consistent with the intended design: a recurrent model can carry its state between steps, while a basic transformer repeatedly processes an expanding context. The comparison should not be read as evidence that PSSA outperforms modern large language models. The developer explicitly says both test models produce poor text at this scale and that larger trials and stronger recurrent baselines are still needed.
Training-speed figures in the documentation need especially careful interpretation. PSSA used GPU acceleration in one run, while the baseline training path was CPU-only. The repository therefore says those headline training rates do not compare the architectures fairly. Its same-CPU measurements show a smaller advantage for PSSA, and the loss comparison is matched by tokens and updates rather than elapsed time.
Several central questions remain open. The project has not yet established whether its reported gap survives models ten or one hundred times larger, whether the memory bank materially improves results, or how the system compares with a modern recurrent baseline. Tests of retention after changing corpora are also unfinished.
The code is publicly available for inspection and accepts issues and pull requests. For now, PSSA is best understood as a transparent experiment with an unusual combination of recurrent state, retrieval-like memory and online adaptation. Its early numbers justify further testing, but broader conclusions depend on independent replication and substantially larger evaluations.



