| Liste des Groupes | Revenir à c theory |
Hi,
Recently there was a paper somebody mentioning
a flit doing a ACK or NACK, to express
backpressure inside a Network on a Chip.
But what is a flit? It seems multiple
flits can be used to create the message
passing in one directiob before the
ACK or NACK in the other direction?
"The growing need for performance from
computing systems drove the industry into
the multi-core and many-core arena. In this
setup, the execution of a kernel (a program)
is split across multiple processors and the
computation happens in parallel
Flits represent logical units of information,
while phits represent the physical domain,
that is, phits represent the number of bits
that can be transferred in parallel in a
single cycle. Consider the Cray T3D. It has
an interconnection network which uses
flit level message flow control wherein each
flit is composed of eight 16-bit phits. That
means its flit size is 128bits and phit size
is 16bits. Also consider the IBM SP2 switch.
It also uses the flit level message flow
control, but its flit size is equal to its
phit size, which is set to 8 bits."
https://en.wikipedia.org/wiki/Flit_(computer_networking)#Example
Well my idea how this is realized in silicon
is rather foggy, I mean even the Hack project
from Nand 2 Tetris, does not show some gate level
schemes for flits and phits.
Could be an interesting extension. But somehow
the image of flits and phits inspired my channel
objects here below. But I am afraid they are fire
and forget, no ACK and NACK:
π-WAM Contest: 1 Million Packets with Prolog
https://medium.com/2989/ec3e91551773
Its amazing that a max_size(1) buffer
can beat an unbounded buffer!
LoL
Bye
Mild Shock schrieb:Hi,
>
How it started, NVIDIA being cool:
>
NCCL provides routines such as all-gather,
all-reduce, broadcast, reduce, reduce-scatter,
and point-to-point send and receive. These
routines are optimized to achieve high
bandwidth and low latency over PCIe,
NVIDIA NVLink™, and other high-speed
interconnects within a node and over
NVIDIA networking across nodes.
https://developer.nvidia.com/nccl
>
How its going, vLLM trying to be cool:
>
[RFC]: Native Weight Syncing APIs
However, there are no standardized methods for
performing online weight syncing. Open source projects
like SkyRL, VeRL, and TRL need to include their
own implementations of the weight syncing
infrastructure, leading to added complexity
for developers seeking to adopt vLLM as their
inference server for post-training workloads.
https://github.com/vllm-project/vllm/issues/31848
>
How much Workers are enough? I guess it depends
on I/O parallelism, CPU Memory parallelism, CPU
Processing parallelism, and now also
>
GPU Memory parallelism and GPU Processing
parallelism, and last but least you might have
a couple DMAs sitting here and there,
>
or even invoking a sort of RDMA. Quite amazing!
>
Bye
>
Mild Shock schrieb:Hi,>
>
Well there are two viewpoint, the "client"
of the GPU, which is the CPU, and the "server"
of the GPU, which is the command processor
>
queue of the GPU device. So basically as
a CPU client I can write the memory area,
that is later mapped to my GPU code storage.
>
And this way have a compiler, even written
in Prolog, that compiles pi-WAM to my Hack VM,
that can then be then deployed to GPU.
>
You could also try the same with a Tiny
LISP VM. And a grown up LISP to act as
the compiler. Would be a similar exercise.
>
Have Fun!
>
Bye
>
Mild Shock schrieb:Hi,>
>
> I've also added comp.theory so Mild Shock can comment.
>
> that has different
> * sizeof ( void * ), and
> * sizeof ( void (*)( void ) ),
>
Could indicate a data RAM and code ROM model.
Which has then the advantage of:
>
Modern operating systems like Windows 11
enforce strict Data Execution Prevention (DEP)
(or NX/XD bit security features) to prevent
malicious programs from injecting and executing
code inside data-only memory regions.
https://root-nation.com/en/soft-en/lifehacks/en-dep-windows-all-about/
>
I adopted data RAM and code ROM model for
pi-WAM from Hack, which has the same separation:
>
Slide 58, Hack Computer
https://drive.google.com/file/d/1Z_fxYmmRNXTkAzmZ6YMoX9NXZIRVCKiw/view
>
But my motivation was not Johnny Depp prevention.
Rather the caching of GPUs. Because WGSL
allows storage annotations read_write and
>
read. I use read_write for the data RAM
of my Hack VM variant, and read for the
code ROM of my Hack VM variant. You can
>
see that here, its open source:
>
@group(0) @binding(0) var<storage, read> code: array<i32>;
@group(0) @binding(1) var<storage, read_write> state: array<i32>;
>
11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget
>
Hope this Helps!
>
Bye
>
Johann 'Myrkraverk' Oskarsson schrieb:On 03/08/2026 6:28 PM, David Brown wrote:>On 03/08/2026 11:41, Richard Harnden wrote:>On 03/08/2026 09:16, David Brown wrote:>>On most targets, function pointers are the same size as void* pointers. But there are exceptions, with some small microcontrollers and DSPs having different kinds of pointers with different sizes, depending on the memory space involved. I have yet to see a situation where there was any reason for storing a function address in a "void*" rather than a more appropriate typedef, such as :>
>
typedef void (*FVoid)(void);
dlsym requires that pointer-to-function is compatible with a void*
>
As I say, I have yet to see a situation where using void* for function pointers was more appropriate than using a function pointer type. If the OS system calls or standard OS libraries makes it a requirement that function pointers are converted to or from void* for some calls, then of course you need to follow those requirements - it's the people who designed the interfaces that made questionable design choices.
>
Nope, you're wrong. You're dead wrong. The world isn't built on C,
even though here in comp.lang.c we like to pretend it is.
>
Several language environments allow function generation on the fly,
these functions need to be garbage collected. Common Lisp is an
example, therefore comp.lang.lisp is added to this discussion.
>
I've also added comp.theory so Mild Shock can comment.
>
You will have to go out of your way to make a computer architecture
incompatible with garbage collected and heap allocated binary code,
something I've been told SBCL does internally [1] to create an archi-
tecture that has different
>
* sizeof ( void * ), and
* sizeof ( void (*)( void ) ),
>
and when you do that, I'll just claim you're making a /malicious
computer architecture/ and refuse to use it.
>
>
[1] I've not looked at the code, but told the garbage collector can
and will at least move the code around, if not collect it.
Les messages affichés proviennent d'usenet.