serbog commited on
Commit
d4c5233
·
1 Parent(s): c099de6

End of training

Browse files
README.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: philschmid/flan-t5-xxl-sharded-fp16
4
+ tags:
5
+ - generated_from_trainer
6
+ model-index:
7
+ - name: temp
8
+ results: []
9
+ library_name: peft
10
+ ---
11
+
12
+ <!-- This model card has been generated automatically according to the information the Trainer had access to. You
13
+ should probably proofread and complete it, then remove this comment. -->
14
+
15
+ # temp
16
+
17
+ This model is a fine-tuned version of [philschmid/flan-t5-xxl-sharded-fp16](https://huggingface.co/philschmid/flan-t5-xxl-sharded-fp16) on the None dataset.
18
+
19
+ ## Model description
20
+
21
+ More information needed
22
+
23
+ ## Intended uses & limitations
24
+
25
+ More information needed
26
+
27
+ ## Training and evaluation data
28
+
29
+ More information needed
30
+
31
+ ## Training procedure
32
+
33
+
34
+ The following `bitsandbytes` quantization config was used during training:
35
+ - load_in_8bit: True
36
+ - load_in_4bit: False
37
+ - llm_int8_threshold: 6.0
38
+ - llm_int8_skip_modules: None
39
+ - llm_int8_enable_fp32_cpu_offload: False
40
+ - llm_int8_has_fp16_weight: False
41
+ - bnb_4bit_quant_type: fp4
42
+ - bnb_4bit_use_double_quant: False
43
+ - bnb_4bit_compute_dtype: float32
44
+ ### Training hyperparameters
45
+
46
+ The following hyperparameters were used during training:
47
+ - learning_rate: 0.001
48
+ - train_batch_size: 8
49
+ - eval_batch_size: 8
50
+ - seed: 42
51
+ - optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
52
+ - lr_scheduler_type: linear
53
+ - num_epochs: 1
54
+
55
+ ### Training results
56
+
57
+ | Training Loss | Epoch | Step | Validation Loss |
58
+ |:-------------:|:-----:|:----:|:---------------:|
59
+ | No log | 1.0 | 266 | 1.7536 |
60
+
61
+
62
+ ### Framework versions
63
+
64
+ - PEFT 0.4.0
65
+ - Transformers 4.31.0
66
+ - Pytorch 2.0.1
67
+ - Datasets 2.14.0
68
+ - Tokenizers 0.13.3
adapter_model.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:f40a2b526d58bb258417d3bae0ec74068606e9029eaef2e671f5241fd026ebd1
3
  size 75604109
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:01b79d0a64a438b6e0ca1accfd49641d69ca7b276bf5ad8c4babc96572426b15
3
  size 75604109
runs/Jul25_01-56-56_2707fad3041e/events.out.tfevents.1690250226.2707fad3041e.4190.8 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:83c5e6c61bd0583a3ccc14db42e7f84bb7194869613842ab548a939e1aae10a5
3
- size 4664
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2666a382e77626677f564f2ded560d72eaa43894e082a425ea67d63b26fd5849
3
+ size 5289