Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on group.lt

3 replies

lemmy.world

I've got a background in deep learning and I still struggle to understand the attention mechanism. I know it's a key/value store but I'm not sure what it's doing to the tensor when it passes through different layers.

4

@behohippy @saint Instead of timestep by timestep sequence modeling the attention allows us to pass sequential model in a parallel NN just like fully connected one, where the positional encoding helps us to know the sequence of each and we can remove the keys having less attention value...

1

You reached the end

The GPT-3 Architecture, on a Napkin | Spyke