Intrinsics for absolute minimum and maximum, and table lookup #324

momchil-velikov · 2024-06-13T16:10:13Z

name: Pull request
about: Technical issues, document format problems, bugs in scripts or feature proposal.

Thank you for submitting a pull request!

If this PR is about a bugfix:

Please use the bugfix label and make sure to go through the checklist below.

If this PR is about a proposal:

We are looking forward to evaluate your proposal, and if possible to
make it part of the Arm C Language Extension (ACLE) specifications.

We would like to encourage you reading through the contribution
guidelines, in particular the section on submitting
a proposal.

Please use the proposal label.

As for any pull request, please make sure to go through the below
checklist.

Checklist: (mark with X those which apply)

If an issue reporting the bug exists, I have mentioned it in the
PR (do not bother creating the issue if all you want to do is
fixing the bug yourself).
I have added/updated the SPDX-FileCopyrightText lines on top
of any file I have edited. Format is SPDX-FileCopyrightText: Copyright {year} {entity or name} <{contact informations}>
(Please update existing copyright lines if applicable. You can
specify year ranges with hyphen , as in 2017-2019, and use
commas to separate gaps, as in 2018-2020, 2022).
I have updated the Copyright section of the sources of the
specification I have edited (this will show up in the text
rendered in the PDF and other output format supported). The
format is the same described in the previous item.
I have run the CI scripts (if applicable, as they might be
tricky to set up on non-*nix machines). The sequence can be
found in the contribution
guidelines. Don't
worry if you cannot run these scripts on your machine, your
patch will be automatically checked in the Actions of the pull
request.
I have added an item that describes the changes I have
introduced in this PR in the section Changes for next
release of the section Change Control/Document history
of the document. Create Changes for next release if it does
not exist. Notice that changes that are not modifying the
content and rendering of the specifications (both HTML and PDF)
do not need to be listed.
When modifying content and/or its rendering, I have checked the
correctness of the result in the PDF output (please refer to the
instructions on how to build the PDFs
locally).
The variable draftversion is set to true in the YAML header
of the sources of the specifications I have modified.
Please DO NOT add my GitHub profile to the list of contributors
in the README page of the project.

neon_intrinsics/advsimd.md

main/acle.md

This patch adds these intrinsics: // Variants are also available for: // [_s8], [_u16], [_s16], [_u32], [_s32], [_u64], [_s64] // [_bf16], [_f16], [_f32], [_f64] void svwrite_lane_zt[_u8](uint64_t zt0, svuint8_t zt, uint64_t idx) __arm_streaming __arm_inout("zt0"); void svwrite_zt[_u8](uint64_t zt0, svuint8_t zt) __arm_streaming __arm_inout("zt0"); according to PR#324[1] [1]ARM-software/acle#324

momchil-velikov · 2024-07-04T17:14:38Z

main/acle.md

+Lookup table read with 4-bit indexes and 8-bit elements.
+``` c
+  // Variants are also available for: _s8
+  svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0");


No, the zn contains indices, it has fixed type, cannot be used for overloading. E.g. the variant for s8 would be

svint8x4_t svluti4_zt_s8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0");

This patch adds these intrinsics: // Variants are also available for: _s8 svuint8x4_t svluti4_zt_u8_x4(uint64_t zt0, svuint8x2_t zn) __arm_streaming __arm_in("zt0"); according to PR#324[1] [1]ARM-software/acle#324

tools/intrinsic_db/advsimd.csv

andrewcarlotti

This patch is missing the SME2 FAMINMAX intrinsics

tools/intrinsic_db/advsimd.csv

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch implements the following intrinsics: * Floating-point absolute maximum (predicated) svfloat16_t svamax[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_z(svbool_t, svfloat16_t, float16_t); * Floating-point absolute minimum (predicated) svfloat16_t svmin[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_z(svbool_t, svfloat16_t, float16_t); All the intrinsics have also variants for `f32` and `f64`, and have the `__arm_streaming` attribute. (cf. ARM-software/acle#324)

This patch implements these intrinsics: ``` c // Variants are also available for: // [_f32_x2], [_f64_x2], // [_f16_x4], [_f32_x4], [_f64_x4] svfloat16x2_t svamax[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; svfloat16x2_t svamin[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; ``` (cf. ARM-software/acle#324) Co-authored-by: Caroline Concatto <[email protected]>

main/acle.md

tools/intrinsic_db/advsimd.csv

rsandifo-arm · 2024-07-31T13:40:13Z

LGTM.

vhscampos · 2024-07-31T14:17:00Z

Thanks for the PR. I can merge it once the conflicts have been resolved.

vhscampos · 2024-08-05T08:20:23Z

There's one little issue in one Copyright header. Please check the build log to spot the problem.

…ntly

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch implements the following intrinsics: * Floating-point absolute maximum (predicated) svfloat16_t svamax[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_z(svbool_t, svfloat16_t, float16_t); * Floating-point absolute minimum (predicated) svfloat16_t svmin[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_z(svbool_t, svfloat16_t, float16_t); All the intrinsics have also variants for `f32` and `f64`, and have the `__arm_streaming` attribute. (cf. ARM-software/acle#324)

momchil-velikov · 2024-09-03T13:05:05Z

Ping?

This patch adds intrinsics for LUTI2 and LUTI4 instructions, which use SVE registers, as specified in the ARM-software/acle#324

This patch adds intrinsics for NEON LUTI2 and LUTI4 instructions as specified in the [ACLE proposal](ARM-software/acle#324)

This patch implements the following intrinsics: * Floating-point absolute maximum (predicated) svfloat16_t svamax[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svamax[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svamax[_n_f16]_z(svbool_t, svfloat16_t, float16_t); * Floating-point absolute minimum (predicated) svfloat16_t svmin[_f16]_m(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_x(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_f16]_z(svbool_t, svfloat16_t, svfloat16_t); svfloat16_t svmin[_n_f16]_m(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_x(svbool_t, svfloat16_t, float16_t); svfloat16_t svmin[_n_f16]_z(svbool_t, svfloat16_t, float16_t); All the intrinsics have also variants for `f32` and `f64`, and have the `__arm_streaming` attribute. (cf. ARM-software/acle#324)

This patch implements these intrinsics: ``` c // Variants are also available for: // [_f32_x2], [_f64_x2], // [_f16_x4], [_f32_x4], [_f64_x4] svfloat16x2_t svamax[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; svfloat16x2_t svamin[_f16_x2](svfloat16x2 zd, svfloat16x2_t zm) __arm_streaming; ``` (cf. ARM-software/acle#324) Co-authored-by: Caroline Concatto <[email protected]>

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 Co-authored-by: Hassnaa Hamdi <[email protected]>

This patch implements the intrinsics of the form floatNxM_t vamin[q]_fN(floatNxM_t vn, floatNxM_t vm); floatNxM_t vamax[q]_fN(floatNxM_t vn, floatNxM_t vm); as defined in ARM-software/acle#324 --------- Co-authored-by: Hassnaa Hamdi <[email protected]>

Lukacma reviewed Jun 21, 2024

View reviewed changes

neon_intrinsics/advsimd.md Outdated Show resolved Hide resolved

Lukacma mentioned this pull request Jun 27, 2024

[AArch64][NEON] Add intrinsics for LUTI llvm/llvm-project#96883

Merged

Lukacma reviewed Jun 27, 2024

View reviewed changes

main/acle.md Show resolved Hide resolved

rsandifo-arm reviewed Jun 28, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

Lukacma mentioned this pull request Jun 28, 2024

[AARCH64][SVE] Add intrinsics for SVE LUTI instructions llvm/llvm-project#97058

Merged

CarolineConcatto reviewed Jul 3, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

CarolineConcatto mentioned this pull request Jul 3, 2024

[Clang][LLVM][AArch64] Add intrinsic for MOVT SME2 instruction llvm/llvm-project#97602

Open

momchil-velikov commented Jul 4, 2024

View reviewed changes

CarolineConcatto mentioned this pull request Jul 4, 2024

[Clang][LLVM][AArch64] Add intrinsic for LUTI4 SME2 instruction llvm/llvm-project#97755

Open

andrewcarlotti reviewed Jul 8, 2024

View reviewed changes

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

andrewcarlotti reviewed Jul 8, 2024

View reviewed changes

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

tools/intrinsic_db/advsimd.csv Outdated Show resolved Hide resolved

This was referenced Jul 16, 2024

[AArch64] Implement NEON vamin/vamax intrinsics llvm/llvm-project#99041

Merged

[AArch64] Implement intrinsics for SVE FAMIN/FAMAX llvm/llvm-project#99042

Merged

momchil-velikov mentioned this pull request Jul 16, 2024

[AArch64] Implement intrinsics for SME2 FAMIN/FAMAX llvm/llvm-project#99063

Merged

ktkachov reviewed Jul 23, 2024

View reviewed changes

main/acle.md Outdated Show resolved Hide resolved

ktkachov reviewed Jul 23, 2024

View reviewed changes

main/acle.md Show resolved Hide resolved

rsandifo-arm reviewed Jul 31, 2024

View reviewed changes

momchil-velikov mentioned this pull request Jul 31, 2024

FP8 ACLE specification #323

Merged

8 tasks

momchil-velikov force-pushed the faminmax-and-luti branch from c4388fa to a6261ca Compare August 1, 2024 09:30

momchil-velikov added 3 commits August 28, 2024 10:27

Intrinsics for absolute minimum and maximum, and table lookup

374eb18

[fixup] Add lane/laneq to some intrinsics, use imm_idx consiste…

d9d0080

…ntly

[fixup] Replace svmovt_zt with svwrite_zt

8a17a84

momchil-velikov added 7 commits August 28, 2024 10:27

[fixup] Add some missing intrinsics, move intrinsics to correct sections

234ebc6

[fixup] Add vluti4q_lane_?8 intrinsics

d25b273

[fixup] Correct a typo

fd6ce52

[fuxup] Add FAMINMAX and LUT to the feature macros table

0f674a5

[fixup] Associate feature test macros with intrinsics

47a942f

[fixup] Misc small fixes

7d77898

[fixup] Update copyright year

e762810

momchil-velikov force-pushed the faminmax-and-luti branch from a6261ca to e762810 Compare August 28, 2024 09:47

vhscampos merged commit e938350 into ARM-software:main Sep 3, 2024
4 checks passed

Lukacma added a commit to llvm/llvm-project that referenced this pull request Sep 4, 2024

[AARCH64][SVE] Add intrinsics for SVE LUTI instructions (#97058)

59093ca

This patch adds intrinsics for LUTI2 and LUTI4 instructions, which use SVE registers, as specified in the ARM-software/acle#324

Lukacma added a commit to llvm/llvm-project that referenced this pull request Sep 4, 2024

[AArch64][NEON] Add intrinsics for LUTI (#96883)

3e948eb

This patch adds intrinsics for NEON LUTI2 and LUTI4 instructions as specified in the [ACLE proposal](ARM-software/acle#324)

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Intrinsics for absolute minimum and maximum, and table lookup #324

Intrinsics for absolute minimum and maximum, and table lookup #324

momchil-velikov commented Jun 13, 2024

momchil-velikov Jul 4, 2024

andrewcarlotti left a comment

rsandifo-arm commented Jul 31, 2024

vhscampos commented Jul 31, 2024

vhscampos commented Aug 5, 2024

momchil-velikov commented Sep 3, 2024

Intrinsics for absolute minimum and maximum, and table lookup #324

Intrinsics for absolute minimum and maximum, and table lookup #324

Conversation

momchil-velikov commented Jun 13, 2024

momchil-velikov Jul 4, 2024

Choose a reason for hiding this comment

andrewcarlotti left a comment

Choose a reason for hiding this comment

rsandifo-arm commented Jul 31, 2024

vhscampos commented Jul 31, 2024

vhscampos commented Aug 5, 2024

momchil-velikov commented Sep 3, 2024