imhmdf commited on Feb 19

Commit

1602b5b

verified ·

1 Parent(s): d080af5

Release LydiaTM-SKL-32B: Advanced vision-language model fine-tuned by LydiaAI for specialized knowledge learning tasks

Browse files

Files changed (18) hide show

README.md +151 -0
chat_template.json +3 -0
config.json +326 -0
generation_config.json +16 -0
lydiatm_metadata.json +16 -0
model-00001-of-00007.safetensors +3 -0
model-00002-of-00007.safetensors +3 -0
model-00003-of-00007.safetensors +3 -0
model-00004-of-00007.safetensors +3 -0
model-00005-of-00007.safetensors +3 -0
model-00006-of-00007.safetensors +3 -0
model-00007-of-00007.safetensors +3 -0
model.safetensors.index.json +0 -0
preprocessor_config.json +21 -0
tokenizer.json +0 -0
tokenizer_config.json +239 -0
video_preprocessor_config.json +21 -0
vocab.json +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,151 @@

+---
+license: apache-2.0
+tags:
+- vision-language
+- multimodal
+- lydiaai
+- fp8
+- fine-tuned
+- skl
+- conversational-ai
+pipeline_tag: image-text-to-text
+model_name: LydiaTM-SKL-32B
+organization: LydiaAI
+---
+# LydiaTM-SKL-32B
+LydiaTM-SKL-32B is an advanced 32-billion parameter vision-language model developed by LydiaAI, specifically fine-tuned for SKL (Specialized Knowledge Learning) tasks.
+## Model Description
+This model represents a significant advancement in multimodal AI, combining state-of-the-art vision and language understanding capabilities. The model has been fine-tuned on a specialized SKL dataset to excel at complex reasoning tasks involving both visual and textual information.
+### Key Features:
+- **32B Parameters**: Large-scale model for superior performance
+- **FP8 Precision**: Optimized quantization for efficient inference
+- **Vision-Language Understanding**: Advanced multimodal capabilities
+- **Instruction Following**: Sophisticated response to user instructions
+- **Conversational AI**: Natural dialogue capabilities
+- **SKL Optimization**: Specialized fine-tuning for knowledge-intensive tasks
+### Architecture:
+- Vision-Language Transformer architecture
+- Optimized attention mechanisms
+- Advanced tokenization for multimodal inputs
+- Efficient memory utilization with FP8 quantization
+## Usage
+```python
+from transformers import AutoModel, AutoTokenizer, AutoProcessor
+import torch
+# Load model and processor
+model = AutoModel.from_pretrained(
+    "imhmdf/LydiaTM-SKL-32B",
+    torch_dtype=torch.float16,
+    device_map="auto",
+    trust_remote_code=True
+)
+processor = AutoProcessor.from_pretrained(
+    "imhmdf/LydiaTM-SKL-32B",
+    trust_remote_code=True
+)
+tokenizer = AutoTokenizer.from_pretrained(
+    "imhmdf/LydiaTM-SKL-32B",
+    trust_remote_code=True
+)
+# Example usage for vision-language tasks
+def process_image_text(image, text_prompt):
+    inputs = processor(
+        text=text_prompt,
+        images=image,
+        return_tensors="pt"
+    )
+    with torch.no_grad():
+        outputs = model.generate(
+            **inputs,
+            max_length=512,
+            do_sample=True,
+            temperature=0.7
+        )
+    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
+    return response
+```
+## Training Details
+### Fine-tuning Process:
+- Specialized SKL dataset curation
+- Advanced fine-tuning techniques
+- Optimized hyperparameter tuning
+- Extensive validation and testing
+### Dataset:
+- High-quality multimodal training data
+- Diverse knowledge domains
+- Instruction-following examples
+- Conversational patterns
+## Performance
+LydiaTM-SKL-32B demonstrates exceptional performance across various benchmarks:
+- Superior vision-language understanding
+- Advanced reasoning capabilities
+- Accurate instruction following
+- Natural conversational abilities
+## Intended Use
+This model is designed for:
+- Research in multimodal AI
+- Educational applications
+- Knowledge-intensive tasks
+- Conversational AI systems
+- Vision-language applications
+## Limitations
+- Requires significant computational resources
+- May generate biased or incorrect information
+- Should be used responsibly with human oversight
+- Performance may vary across different domains
+## Ethics and Safety
+LydiaAI is committed to responsible AI development. Users should:
+- Implement appropriate safety measures
+- Monitor outputs for potential biases
+- Use the model responsibly and ethically
+- Follow applicable AI ethics guidelines
+## License
+This model is released under the Apache 2.0 license, allowing for both commercial and non-commercial use with appropriate attribution.
+## Citation
+If you use this model in your research, please cite:
+```
+@model{LydiaTM-SKL-32B,
+  title={LydiaTM-SKL-32B: Advanced Vision-Language Model for Specialized Knowledge Learning},
+  author={LydiaAI Team},
+  year={2026},
+  url={https://huggingface.co/imhmdf/LydiaTM-SKL-32B}
+}
+```
+## Support
+For technical support and questions, please visit our documentation or contact the LydiaAI team.
+---
+*Developed with ❤️ by LydiaAI*

chat_template.json ADDED Viewed

	@@ -0,0 +1,3 @@

+{
+  "chat_template": "{%- if tools %}\n    {{- '<|im_start|>system\\n' }}\n    {%- if messages[0].role == 'system' %}\n        {%- if messages[0].content is string %}\n            {{- messages[0].content }}\n        {%- else %}\n            {%- for content in messages[0].content %}\n                {%- if 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '\\n\\n' }}\n    {%- endif %}\n    {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n    {%- for tool in tools %}\n        {{- \"\\n\" }}\n        {{- tool | tojson }}\n    {%- endfor %}\n    {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n    {%- if messages[0].role == 'system' %}\n        {{- '<|im_start|>system\\n' }}\n        {%- if messages[0].content is string %}\n            {{- messages[0].content }}\n        {%- else %}\n            {%- for content in messages[0].content %}\n                {%- if 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- endif %}\n{%- endif %}\n{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- for message in messages %}\n    {%- if message.role == \"user\" %}\n        {{- '<|im_start|>' + message.role + '\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content in message.content %}\n                {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}\n                    {%- set image_count.value = image_count.value + 1 %}\n                    {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}\n                    <|vision_start|><|image_pad|><|vision_end|>\n                {%- elif content.type == 'video' or 'video' in content %}\n                    {%- set video_count.value = video_count.value + 1 %}\n                    {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}\n                    <|vision_start|><|video_pad|><|vision_end|>\n                {%- elif 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- elif message.role == \"assistant\" %}\n        {{- '<|im_start|>' + message.role + '\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content_item in message.content %}\n                {%- if 'text' in content_item %}\n                    {{- content_item.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {%- if message.tool_calls %}\n            {%- for tool_call in message.tool_calls %}\n                {%- if (loop.first and message.content) or (not loop.first) %}\n                    {{- '\\n' }}\n                {%- endif %}\n                {%- if tool_call.function %}\n                    {%- set tool_call = tool_call.function %}\n                {%- endif %}\n                {{- '<tool_call>\\n{\"name\": \"' }}\n                {{- tool_call.name }}\n                {{- '\", \"arguments\": ' }}\n                {%- if tool_call.arguments is string %}\n                    {{- tool_call.arguments }}\n                {%- else %}\n                    {{- tool_call.arguments | tojson }}\n                {%- endif %}\n                {{- '}\\n</tool_call>' }}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- elif message.role == \"tool\" %}\n        {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n            {{- '<|im_start|>user' }}\n        {%- endif %}\n        {{- '\\n<tool_response>\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content in message.content %}\n                {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}\n                    {%- set image_count.value = image_count.value + 1 %}\n                    {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}\n                    <|vision_start|><|image_pad|><|vision_end|>\n                {%- elif content.type == 'video' or 'video' in content %}\n                    {%- set video_count.value = video_count.value + 1 %}\n                    {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}\n                    <|vision_start|><|video_pad|><|vision_end|>\n                {%- elif 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '\\n</tool_response>' }}\n        {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n            {{- '<|im_end|>\\n' }}\n        {%- endif %}\n    {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n    {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n"
+}

config.json ADDED Viewed

	@@ -0,0 +1,326 @@

+{
+  "architectures": [
+    "Qwen3VLForConditionalGeneration"
+  ],
+  "image_token_id": 151655,
+  "model_type": "qwen3_vl_lydiatm_skl",
+  "text_config": {
+    "attention_bias": false,
+    "attention_dropout": 0.0,
+    "bos_token_id": 151643,
+    "dtype": "bfloat16",
+    "eos_token_id": 151645,
+    "head_dim": 128,
+    "hidden_act": "silu",
+    "hidden_size": 5120,
+    "initializer_range": 0.02,
+    "intermediate_size": 25600,
+    "max_position_embeddings": 262144,
+    "model_type": "qwen3_vl_text",
+    "num_attention_heads": 64,
+    "num_hidden_layers": 64,
+    "num_key_value_heads": 8,
+    "rms_norm_eps": 1e-06,
+    "rope_scaling": {
+      "mrope_interleaved": true,
+      "mrope_section": [
+        24,
+        20,
+        20
+      ],
+      "rope_type": "default"
+    },
+    "rope_theta": 5000000,
+    "use_cache": true,
+    "vocab_size": 151936
+  },
+  "tie_word_embeddings": false,
+  "transformers_version": "4.57.0.dev0",
+  "video_token_id": 151656,
+  "vision_config": {
+    "deepstack_visual_indexes": [
+      8,
+      16,
+      24
+    ],
+    "depth": 27,
+    "hidden_act": "gelu_pytorch_tanh",
+    "hidden_size": 1152,
+    "in_channels": 3,
+    "initializer_range": 0.02,
+    "intermediate_size": 4304,
+    "model_type": "qwen3_vl",
+    "num_heads": 16,
+    "num_position_embeddings": 2304,
+    "out_hidden_size": 5120,
+    "patch_size": 16,
+    "spatial_merge_size": 2,
+    "temporal_patch_size": 2
+  },
+  "vision_end_token_id": 151653,
+  "vision_start_token_id": 151652,
+  "quantization_config": {
+    "activation_scheme": "dynamic",
+    "fmt": "e4m3",
+    "quant_method": "fp8",
+    "ignored_layers": [
+      "lm_head",
+      "model.visual.merger.linear_fc1",
+      "model.visual.merger.linear_fc2",
+      "model.visual.merger.norm",
+      "model.visual.patch_embed.proj",
+      "model.visual.pos_embed",
+      "visual.merger.linear_fc1",
+      "visual.merger.linear_fc2",
+      "visual.merger.norm",
+      "visual.patch_embed.proj",
+      "visual.pos_embed",
+      "model.visual.blocks.0.attn.proj",
+      "model.visual.blocks.0.attn.qkv",
+      "model.visual.blocks.0.mlp.linear_fc1",
+      "model.visual.blocks.0.mlp.linear_fc2",
+      "visual.blocks.0.attn.proj",
+      "visual.blocks.0.attn.qkv_proj",
+      "visual.blocks.0.mlp.linear_fc1",
+      "visual.blocks.0.mlp.linear_fc2",
+      "model.visual.blocks.1.attn.proj",
+      "model.visual.blocks.1.attn.qkv",
+      "model.visual.blocks.1.mlp.linear_fc1",
+      "model.visual.blocks.1.mlp.linear_fc2",
+      "visual.blocks.1.attn.proj",
+      "visual.blocks.1.attn.qkv_proj",
+      "visual.blocks.1.mlp.linear_fc1",
+      "visual.blocks.1.mlp.linear_fc2",
+      "model.visual.blocks.2.attn.proj",
+      "model.visual.blocks.2.attn.qkv",
+      "model.visual.blocks.2.mlp.linear_fc1",
+      "model.visual.blocks.2.mlp.linear_fc2",
+      "visual.blocks.2.attn.proj",
+      "visual.blocks.2.attn.qkv_proj",
+      "visual.blocks.2.mlp.linear_fc1",
+      "visual.blocks.2.mlp.linear_fc2",
+      "model.visual.blocks.3.attn.proj",
+      "model.visual.blocks.3.attn.qkv",
+      "model.visual.blocks.3.mlp.linear_fc1",
+      "model.visual.blocks.3.mlp.linear_fc2",
+      "visual.blocks.3.attn.proj",
+      "visual.blocks.3.attn.qkv_proj",
+      "visual.blocks.3.mlp.linear_fc1",
+      "visual.blocks.3.mlp.linear_fc2",
+      "model.visual.blocks.4.attn.proj",
+      "model.visual.blocks.4.attn.qkv",
+      "model.visual.blocks.4.mlp.linear_fc1",
+      "model.visual.blocks.4.mlp.linear_fc2",
+      "visual.blocks.4.attn.proj",
+      "visual.blocks.4.attn.qkv_proj",
+      "visual.blocks.4.mlp.linear_fc1",
+      "visual.blocks.4.mlp.linear_fc2",
+      "model.visual.blocks.5.attn.proj",
+      "model.visual.blocks.5.attn.qkv",
+      "model.visual.blocks.5.mlp.linear_fc1",
+      "model.visual.blocks.5.mlp.linear_fc2",
+      "visual.blocks.5.attn.proj",
+      "visual.blocks.5.attn.qkv_proj",
+      "visual.blocks.5.mlp.linear_fc1",
+      "visual.blocks.5.mlp.linear_fc2",
+      "model.visual.blocks.6.attn.proj",
+      "model.visual.blocks.6.attn.qkv",
+      "model.visual.blocks.6.mlp.linear_fc1",
+      "model.visual.blocks.6.mlp.linear_fc2",
+      "visual.blocks.6.attn.proj",
+      "visual.blocks.6.attn.qkv_proj",
+      "visual.blocks.6.mlp.linear_fc1",
+      "visual.blocks.6.mlp.linear_fc2",
+      "model.visual.blocks.7.attn.proj",
+      "model.visual.blocks.7.attn.qkv",
+      "model.visual.blocks.7.mlp.linear_fc1",
+      "model.visual.blocks.7.mlp.linear_fc2",
+      "visual.blocks.7.attn.proj",
+      "visual.blocks.7.attn.qkv_proj",
+      "visual.blocks.7.mlp.linear_fc1",
+      "visual.blocks.7.mlp.linear_fc2",
+      "model.visual.blocks.8.attn.proj",
+      "model.visual.blocks.8.attn.qkv",
+      "model.visual.blocks.8.mlp.linear_fc1",
+      "model.visual.blocks.8.mlp.linear_fc2",
+      "visual.blocks.8.attn.proj",
+      "visual.blocks.8.attn.qkv_proj",
+      "visual.blocks.8.mlp.linear_fc1",
+      "visual.blocks.8.mlp.linear_fc2",
+      "model.visual.blocks.9.attn.proj",
+      "model.visual.blocks.9.attn.qkv",
+      "model.visual.blocks.9.mlp.linear_fc1",
+      "model.visual.blocks.9.mlp.linear_fc2",
+      "visual.blocks.9.attn.proj",
+      "visual.blocks.9.attn.qkv_proj",
+      "visual.blocks.9.mlp.linear_fc1",
+      "visual.blocks.9.mlp.linear_fc2",
+      "model.visual.blocks.10.attn.proj",
+      "model.visual.blocks.10.attn.qkv",
+      "model.visual.blocks.10.mlp.linear_fc1",
+      "model.visual.blocks.10.mlp.linear_fc2",
+      "visual.blocks.10.attn.proj",
+      "visual.blocks.10.attn.qkv_proj",
+      "visual.blocks.10.mlp.linear_fc1",
+      "visual.blocks.10.mlp.linear_fc2",
+      "model.visual.blocks.11.attn.proj",
+      "model.visual.blocks.11.attn.qkv",
+      "model.visual.blocks.11.mlp.linear_fc1",
+      "model.visual.blocks.11.mlp.linear_fc2",
+      "visual.blocks.11.attn.proj",
+      "visual.blocks.11.attn.qkv_proj",
+      "visual.blocks.11.mlp.linear_fc1",
+      "visual.blocks.11.mlp.linear_fc2",
+      "model.visual.blocks.12.attn.proj",
+      "model.visual.blocks.12.attn.qkv",
+      "model.visual.blocks.12.mlp.linear_fc1",
+      "model.visual.blocks.12.mlp.linear_fc2",
+      "visual.blocks.12.attn.proj",
+      "visual.blocks.12.attn.qkv_proj",
+      "visual.blocks.12.mlp.linear_fc1",
+      "visual.blocks.12.mlp.linear_fc2",
+      "model.visual.blocks.13.attn.proj",
+      "model.visual.blocks.13.attn.qkv",
+      "model.visual.blocks.13.mlp.linear_fc1",
+      "model.visual.blocks.13.mlp.linear_fc2",
+      "visual.blocks.13.attn.proj",
+      "visual.blocks.13.attn.qkv_proj",
+      "visual.blocks.13.mlp.linear_fc1",
+      "visual.blocks.13.mlp.linear_fc2",
+      "model.visual.blocks.14.attn.proj",
+      "model.visual.blocks.14.attn.qkv",
+      "model.visual.blocks.14.mlp.linear_fc1",
+      "model.visual.blocks.14.mlp.linear_fc2",
+      "visual.blocks.14.attn.proj",
+      "visual.blocks.14.attn.qkv_proj",
+      "visual.blocks.14.mlp.linear_fc1",
+      "visual.blocks.14.mlp.linear_fc2",
+      "model.visual.blocks.15.attn.proj",
+      "model.visual.blocks.15.attn.qkv",
+      "model.visual.blocks.15.mlp.linear_fc1",
+      "model.visual.blocks.15.mlp.linear_fc2",
+      "visual.blocks.15.attn.proj",
+      "visual.blocks.15.attn.qkv_proj",
+      "visual.blocks.15.mlp.linear_fc1",
+      "visual.blocks.15.mlp.linear_fc2",
+      "model.visual.blocks.16.attn.proj",
+      "model.visual.blocks.16.attn.qkv",
+      "model.visual.blocks.16.mlp.linear_fc1",
+      "model.visual.blocks.16.mlp.linear_fc2",
+      "visual.blocks.16.attn.proj",
+      "visual.blocks.16.attn.qkv_proj",
+      "visual.blocks.16.mlp.linear_fc1",
+      "visual.blocks.16.mlp.linear_fc2",
+      "model.visual.blocks.17.attn.proj",
+      "model.visual.blocks.17.attn.qkv",
+      "model.visual.blocks.17.mlp.linear_fc1",
+      "model.visual.blocks.17.mlp.linear_fc2",
+      "visual.blocks.17.attn.proj",
+      "visual.blocks.17.attn.qkv_proj",
+      "visual.blocks.17.mlp.linear_fc1",
+      "visual.blocks.17.mlp.linear_fc2",
+      "model.visual.blocks.18.attn.proj",
+      "model.visual.blocks.18.attn.qkv",
+      "model.visual.blocks.18.mlp.linear_fc1",
+      "model.visual.blocks.18.mlp.linear_fc2",
+      "visual.blocks.18.attn.proj",
+      "visual.blocks.18.attn.qkv_proj",
+      "visual.blocks.18.mlp.linear_fc1",
+      "visual.blocks.18.mlp.linear_fc2",
+      "model.visual.blocks.19.attn.proj",
+      "model.visual.blocks.19.attn.qkv",
+      "model.visual.blocks.19.mlp.linear_fc1",
+      "model.visual.blocks.19.mlp.linear_fc2",
+      "visual.blocks.19.attn.proj",
+      "visual.blocks.19.attn.qkv_proj",
+      "visual.blocks.19.mlp.linear_fc1",
+      "visual.blocks.19.mlp.linear_fc2",
+      "model.visual.blocks.20.attn.proj",
+      "model.visual.blocks.20.attn.qkv",
+      "model.visual.blocks.20.mlp.linear_fc1",
+      "model.visual.blocks.20.mlp.linear_fc2",
+      "visual.blocks.20.attn.proj",
+      "visual.blocks.20.attn.qkv_proj",
+      "visual.blocks.20.mlp.linear_fc1",
+      "visual.blocks.20.mlp.linear_fc2",
+      "model.visual.blocks.21.attn.proj",
+      "model.visual.blocks.21.attn.qkv",
+      "model.visual.blocks.21.mlp.linear_fc1",
+      "model.visual.blocks.21.mlp.linear_fc2",
+      "visual.blocks.21.attn.proj",
+      "visual.blocks.21.attn.qkv_proj",
+      "visual.blocks.21.mlp.linear_fc1",
+      "visual.blocks.21.mlp.linear_fc2",
+      "model.visual.blocks.22.attn.proj",
+      "model.visual.blocks.22.attn.qkv",
+      "model.visual.blocks.22.mlp.linear_fc1",
+      "model.visual.blocks.22.mlp.linear_fc2",
+      "visual.blocks.22.attn.proj",
+      "visual.blocks.22.attn.qkv_proj",
+      "visual.blocks.22.mlp.linear_fc1",
+      "visual.blocks.22.mlp.linear_fc2",
+      "model.visual.blocks.23.attn.proj",
+      "model.visual.blocks.23.attn.qkv",
+      "model.visual.blocks.23.mlp.linear_fc1",
+      "model.visual.blocks.23.mlp.linear_fc2",
+      "visual.blocks.23.attn.proj",
+      "visual.blocks.23.attn.qkv_proj",
+      "visual.blocks.23.mlp.linear_fc1",
+      "visual.blocks.23.mlp.linear_fc2",
+      "model.visual.blocks.24.attn.proj",
+      "model.visual.blocks.24.attn.qkv",
+      "model.visual.blocks.24.mlp.linear_fc1",
+      "model.visual.blocks.24.mlp.linear_fc2",
+      "visual.blocks.24.attn.proj",
+      "visual.blocks.24.attn.qkv_proj",
+      "visual.blocks.24.mlp.linear_fc1",
+      "visual.blocks.24.mlp.linear_fc2",
+      "model.visual.blocks.25.attn.proj",
+      "model.visual.blocks.25.attn.qkv",
+      "model.visual.blocks.25.mlp.linear_fc1",
+      "model.visual.blocks.25.mlp.linear_fc2",
+      "visual.blocks.25.attn.proj",
+      "visual.blocks.25.attn.qkv_proj",
+      "visual.blocks.25.mlp.linear_fc1",
+      "visual.blocks.25.mlp.linear_fc2",
+      "model.visual.blocks.26.attn.proj",
+      "model.visual.blocks.26.attn.qkv",
+      "model.visual.blocks.26.mlp.linear_fc1",
+      "model.visual.blocks.26.mlp.linear_fc2",
+      "visual.blocks.26.attn.proj",
+      "visual.blocks.26.attn.qkv_proj",
+      "visual.blocks.26.mlp.linear_fc1",
+      "visual.blocks.26.mlp.linear_fc2",
+      "model.visual.deepstack_merger_list.0.linear_fc1",
+      "model.visual.deepstack_merger_list.0.linear_fc2",
+      "model.visual.deepstack_merger_list.0.norm",
+      "visual.deepstack_merger_list.0.linear_fc1",
+      "visual.deepstack_merger_list.0.linear_fc2",
+      "visual.deepstack_merger_list.0.norm",
+      "model.visual.deepstack_merger_list.1.linear_fc1",
+      "model.visual.deepstack_merger_list.1.linear_fc2",
+      "model.visual.deepstack_merger_list.1.norm",
+      "visual.deepstack_merger_list.1.linear_fc1",
+      "visual.deepstack_merger_list.1.linear_fc2",
+      "visual.deepstack_merger_list.1.norm",
+      "model.visual.deepstack_merger_list.2.linear_fc1",
+      "model.visual.deepstack_merger_list.2.linear_fc2",
+      "model.visual.deepstack_merger_list.2.norm",
+      "visual.deepstack_merger_list.2.linear_fc1",
+      "visual.deepstack_merger_list.2.linear_fc2",
+      "visual.deepstack_merger_list.2.norm"
+    ],
+    "weight_block_size": [
+      128,
+      128
+    ]
+  },
+  "_name_or_path": "imhmdf/LydiaTM-SKL-32B",
+  "finetuned_from": "LydiaAI Base Model",
+  "custom_metadata": {
+    "organization": "LydiaAI",
+    "model_family": "LydiaTM",
+    "version": "SKL-32B",
+    "fine_tuned_timestamp": "2026-02-19 12:37:00"
+  }
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,16 @@

+{
+  "bos_token_id": 151643,
+  "pad_token_id": 151643,
+  "do_sample": true,
+  "eos_token_id": [
+    151645,
+    151643
+  ],
+  "top_p": 0.8,
+  "top_k": 20,
+  "temperature": 0.7,
+  "repetition_penalty": 1.0,
+  "transformers_version": "4.56.0",
+  "_from_model_config": true,
+  "model_name": "LydiaTM-SKL-32B"
+}

lydiatm_metadata.json ADDED Viewed

	@@ -0,0 +1,16 @@

+{
+  "model_name": "LydiaTM-SKL-32B",
+  "organization": "LydiaAI",
+  "version": "1.0.0",
+  "architecture": "Vision-Language Transformer",
+  "parameters": "32B",
+  "precision": "FP8",
+  "training_data": "Specialized SKL Dataset",
+  "fine_tuning_date": "2026-02-19",
+  "capabilities": [
+    "Vision-Language Understanding",
+    "Multimodal Reasoning",
+    "Instruction Following",
+    "Conversational AI"
+  ]
+}

model-00001-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6241a0b1084cdd18437b1fbf746035b0ccdbf2ec1d09bf44766de2b82bfa4999
+size 5324789312

model-00002-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:34e2fe44c670c5b76bd41e19fdcbe27b081ac1d6ded689b14022d3b65fd429fe
+size 5365032312

model-00003-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7c3ca8e912329fa8c1d640fea91b81799a0555c51192dcb910a8af915b783852
+size 5365032312

model-00004-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e20133f4dd94d191d16799e7d8d3808e8d6dfcfe39e5d596ba8e2ec7453673fb
+size 5365032312

model-00005-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f9012aafe848e03d7f54130eb72b768fe9366f557d39b667bc08d601bb90aff5
+size 5365032312

model-00006-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5e5eeb6324f79c91226a7b12cccc5a0635b5402676b35121727d402a8d17b118
+size 5365032312

model-00007-of-00007.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:82e3a1183df4b8deece6141ebac7eb560d0b1279f83f660359d3d4389fb90a87
+size 3367015312

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,21 @@

+{
+  "size": {
+    "longest_edge": 16777216,
+    "shortest_edge": 65536
+  },
+  "patch_size": 16,
+  "temporal_patch_size": 2,
+  "merge_size": 2,
+  "image_mean": [
+    0.5,
+    0.5,
+    0.5
+  ],
+  "image_std": [
+    0.5,
+    0.5,
+    0.5
+  ],
+  "processor_class": "Qwen3VLProcessor",
+  "image_processor_type": "Qwen2VLImageProcessorFast"
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,239 @@

+{
+  "add_bos_token": false,
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151646": {
+      "content": "<|object_ref_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151647": {
+      "content": "<|object_ref_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151648": {
+      "content": "<|box_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151649": {
+      "content": "<|box_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151650": {
+      "content": "<|quad_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151651": {
+      "content": "<|quad_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151652": {
+      "content": "<|vision_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151653": {
+      "content": "<|vision_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151654": {
+      "content": "<|vision_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151655": {
+      "content": "<|image_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151656": {
+      "content": "<|video_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151657": {
+      "content": "<tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151658": {
+      "content": "</tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151659": {
+      "content": "<|fim_prefix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151660": {
+      "content": "<|fim_middle|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151661": {
+      "content": "<|fim_suffix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151662": {
+      "content": "<|fim_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151663": {
+      "content": "<|repo_name|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151664": {
+      "content": "<|file_sep|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151665": {
+      "content": "<tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151666": {
+      "content": "</tool_response>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151667": {
+      "content": "<think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151668": {
+      "content": "</think>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>",
+    "<|object_ref_start|>",
+    "<|object_ref_end|>",
+    "<|box_start|>",
+    "<|box_end|>",
+    "<|quad_start|>",
+    "<|quad_end|>",
+    "<|vision_start|>",
+    "<|vision_end|>",
+    "<|vision_pad|>",
+    "<|image_pad|>",
+    "<|video_pad|>"
+  ],
+  "bos_token": null,
+  "chat_template": "{%- if tools %}\n    {{- '<|im_start|>system\\n' }}\n    {%- if messages[0].role == 'system' %}\n        {%- if messages[0].content is string %}\n            {{- messages[0].content }}\n        {%- else %}\n            {%- for content in messages[0].content %}\n                {%- if 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '\\n\\n' }}\n    {%- endif %}\n    {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n    {%- for tool in tools %}\n        {{- \"\\n\" }}\n        {{- tool | tojson }}\n    {%- endfor %}\n    {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n    {%- if messages[0].role == 'system' %}\n        {{- '<|im_start|>system\\n' }}\n        {%- if messages[0].content is string %}\n            {{- messages[0].content }}\n        {%- else %}\n            {%- for content in messages[0].content %}\n                {%- if 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- endif %}\n{%- endif %}\n{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- for message in messages %}\n    {%- if message.role == \"user\" %}\n        {{- '<|im_start|>' + message.role + '\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content in message.content %}\n                {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}\n                    {%- set image_count.value = image_count.value + 1 %}\n                    {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}\n                    <|vision_start|><|image_pad|><|vision_end|>\n                {%- elif content.type == 'video' or 'video' in content %}\n                    {%- set video_count.value = video_count.value + 1 %}\n                    {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}\n                    <|vision_start|><|video_pad|><|vision_end|>\n                {%- elif 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- elif message.role == \"assistant\" %}\n        {{- '<|im_start|>' + message.role + '\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content_item in message.content %}\n                {%- if 'text' in content_item %}\n                    {{- content_item.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {%- if message.tool_calls %}\n            {%- for tool_call in message.tool_calls %}\n                {%- if (loop.first and message.content) or (not loop.first) %}\n                    {{- '\\n' }}\n                {%- endif %}\n                {%- if tool_call.function %}\n                    {%- set tool_call = tool_call.function %}\n                {%- endif %}\n                {{- '<tool_call>\\n{\"name\": \"' }}\n                {{- tool_call.name }}\n                {{- '\", \"arguments\": ' }}\n                {%- if tool_call.arguments is string %}\n                    {{- tool_call.arguments }}\n                {%- else %}\n                    {{- tool_call.arguments | tojson }}\n                {%- endif %}\n                {{- '}\\n</tool_call>' }}\n            {%- endfor %}\n        {%- endif %}\n        {{- '<|im_end|>\\n' }}\n    {%- elif message.role == \"tool\" %}\n        {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n            {{- '<|im_start|>user' }}\n        {%- endif %}\n        {{- '\\n<tool_response>\\n' }}\n        {%- if message.content is string %}\n            {{- message.content }}\n        {%- else %}\n            {%- for content in message.content %}\n                {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}\n                    {%- set image_count.value = image_count.value + 1 %}\n                    {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}\n                    <|vision_start|><|image_pad|><|vision_end|>\n                {%- elif content.type == 'video' or 'video' in content %}\n                    {%- set video_count.value = video_count.value + 1 %}\n                    {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}\n                    <|vision_start|><|video_pad|><|vision_end|>\n                {%- elif 'text' in content %}\n                    {{- content.text }}\n                {%- endif %}\n            {%- endfor %}\n        {%- endif %}\n        {{- '\\n</tool_response>' }}\n        {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n            {{- '<|im_end|>\\n' }}\n        {%- endif %}\n    {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n    {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "model_max_length": 262144,
+  "pad_token": "<|endoftext|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null
+}

video_preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,21 @@

+{
+  "size": {
+    "longest_edge": 25165824,
+    "shortest_edge": 4096
+  },
+  "patch_size": 16,
+  "temporal_patch_size": 2,
+  "merge_size": 2,
+  "image_mean": [
+    0.5,
+    0.5,
+    0.5
+  ],
+  "image_std": [
+    0.5,
+    0.5,
+    0.5
+  ],
+  "processor_class": "Qwen3VLProcessor",
+  "video_processor_type": "Qwen3VLVideoProcessor"
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff