Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on lemmy.fmhy.ml
singularity·Singularity | Artificial Intelligence (ai), Technology & FuturologybyMartineski

Derivative Free Weight-space Ensembling (paper from 7.07.2023)

Recent work suggests that interpolating between the weights of two specialized language models can transfer knowledge between tasks in a way that multi-task learning cannot. However, very few have explored interpolation between more than two models, where each has a distinct knowledge base. In this paper, we introduce Derivative Free Weight-space Ensembling (DFWE), a new few-sample task transfer approach for open-domain dialogue. Our framework creates a set of diverse expert language models trained using a predefined set of source tasks. Next, we finetune each of the expert models on the target task, approaching the target task from several distinct knowledge bases. Finally, we linearly interpolate between the model weights using a gradient-free-optimization algorithm, to efficiently find a good interpolation weighting. We demonstrate the effectiveness of the method on FETA-Friends outperforming the standard pretrain-finetune approach.

https://arxiv.org/abs/2307.03506Open linkView original on lemmy.fmhy.ml
2

No replies yet

No comments on the original post yet.
Derivative Free Weight-space Ensembling (paper from 7.07.2023) | Spyke