what pretraining looks like #CATCHAT is using a stat-of-the-art GRU-L3-V3a #LLM written with #LLVM #CPP and #Vulkan #compute #shaders . no python, training on our first #dataset of dialogue 397M dialogue turns tokens and 78.9m parameters as for our phase0 #pretraining running on 1x L40S by #NVIDIA
