gwishy-deep-pro / README.md
MOTU01's picture
Gwishy Deep check Pro: flan-t5-large-grammar-synthesis ONNX (q4 weight-only + int8)
04ca191 verified
|
Raw
History Blame Contribute Delete
1.95 kB
metadata
license: apache-2.0
base_model: pszemraj/flan-t5-large-grammar-synthesis
pipeline_tag: text2text-generation
library_name: transformers.js
tags:
  - grammar
  - grammar-correction
  - onnx
  - transformers.js
  - gwishy

Gwishy Deep Check Pro — flan-t5-large-grammar-synthesis (ONNX, 4-bit)

ONNX conversion of pszemraj/flan-t5-large-grammar-synthesis (783M parameters) packaged for transformers.js, built for the Deep check Pro tier of the Gwishy writing assistant.

  • onnx/*_q4.onnxMatMul 4-bit weight-only (block 64, symmetric). Activations stay fp32, so correction quality holds where naive dynamic int8 quantization of this model family hallucinates. ~688 MB total. This is what Gwishy loads (dtype: "q4").
  • onnx/*_int8.onnx — dynamic int8, included for comparison (~819 MB).

Conversion pipeline: optimum-cli export onnx (task text2text-generation-with-past), decoder merge via optimum.onnx.merge_decoders, quantization via onnxruntime MatMulNBitsQuantizer — see scripts/convert-pro-model.py.

Usage (transformers.js)

import { pipeline } from '@huggingface/transformers';
const fix = await pipeline('text2text-generation', 'MOTU01/gwishy-deep-pro', { dtype: 'q4' });
const out = await fix('Furthermore, the teams research proposal was deferred.', { max_new_tokens: 192 });
// -> "Furthermore, the team's research proposal was deferred."

License

The base model is dual-licensed by its author (Apache-2.0 / CC-BY-NC-SA-4.0); this conversion is distributed under the Apache-2.0 option. All credit for the model itself to @pszemraj.