ArtifartX@alien.topB to

LocalLLaMAEnglish · 2 years ago

Point me towards some basic dataset preparation tips for LLM's?

1

Point me towards some basic dataset preparation tips for LLM's?

ArtifartX@alien.topB to

LocalLLaMAEnglish · 2 years ago

I have some basic confusions over how to prepare a dataset for training. My plan is to use a model like llama2 7b chat, and train it on some proprietary data I have (in its raw format, this data is very similar to a text book). Do I need to find a way to reformat this large amount of text into a bunch of pairs like “query” and “output” ?

I have seen some LLM’s which say things like “trained on Wikipedia” which seems like they were able to train it on that large chunk of text alone without reformatting it into data pairs - is there a way I can do that, too? Or since I want to target a chat model, I have to find a way to convert the data into pairs which basically serve as examples of proper input and output?

Chat

ArtifartX@alien.topOPB
link
fedilink
English
arrow-up
1·
2 years ago
Awesome, thank you, at a glance this looks like it will be very helpful