| Age | Commit message (Collapse) | Author |
|
Some configs with gcc are also now affected.
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why&How]
When building the DML files with clang without any sanitizer or LTO,
the following -Wframe-larger-than errors break the build under
CONFIG_WERROR:
display_mode_vba_30.c: error: stack frame size (2512) exceeds limit
(2048) in 'dml30_ModeSupportAndSystemConfigurationFull'
display_mode_vba_31.c: error: stack frame size (2416) exceeds limit
(2048) in 'dml31_ModeSupportAndSystemConfigurationFull'
display_mode_vba_314.c: error: stack frame size (2392) exceeds limit
(2048) in 'dml314_ModeSupportAndSystemConfigurationFull'
Clang consistently spills more than gcc, pushing the frame past the 2048
byte limit.
Apply an existing approach of increasing the warn stack size to the
non-sanitizer path so plain clang builds use a 3072 byte limit.
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[WHY]
UBSAN instrumentation adds checks and handler calls and increases
stack usage in the large DML calculation functions, similar to
KASAN and KCSAN. With UBSAN enabled these files exceed the default
-Wframe-larger-than limit and fail to build when -Werror is in effect.
Reproduced with LLVM (make LLVM=1, clang 19.1.1), CONFIG_UBSAN=y,
CONFIG_GCOV_PROFILE_ALL=y and CONFIG_DRM_AMDGPU_WERROR=y on x86_64.
[HOW]
Include CONFIG_UBSAN in the sanitizer check that selects the
higher per-file frame warning limit in the dml and dml2_0
Makefiles.
Suggested-by: Leo Li <sunpeng.li@amd.com>
Assisted-by: Copilot:Claude-Opus-5.5
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
Use of compound literals to initialize heap objects may create stack
temporaries when compiling in debug mode. For large objects such
as dc_memory_pool which is 2 pages large to avoid cacheline sharing,
this causes function stackframe to exceed kernel limit.
[How]
Replace compound literal with field-by-field initialization.
Reported-by: Mark Brown <broonie@kernel.org>
Link: https://lore.kernel.org/all/aq0vYhZXtMqA41wZ@sirena.org.uk/
Signed-off-by: Dominik Kaszewski <dominik.kaszewski@amd.com>
Signed-off-by: Leo Li <sunpeng.li@amd.com>
Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Reviewed-by: Aric Cyr <aric.cyr@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Smatch reports that link->link_enc is unconditionally dereferenced at
the top of dce110_enable_tmds_link_output() while the later
setup_ri_pj_check_in_sw_or_hw_mode calls guard against it being NULL,
making those checks contradictory.
Add an early return when link_enc is NULL and drop the redundant NULL
checks from the two call sites that follow, since link_enc is guaranteed
non-NULL past the guard
Fixes: 992694ad2858 ("drm/amd/display: Enable DCN6 sources compilation")
Reported-by: Dan Carpenter <error27@gmail.com>
Cc: Aurabindo Pillai <aurabindo.pillai@amd.com>
Cc: Roman Li <Roman.Li@amd.com>
Cc: Ivan Lipski <ivan.lipski@amd.com>
Cc: Dan Wheeler <daniel.wheeler@amd.com>
Cc: Alex Hung <alex.hung@amd.com>
Cc: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Reviewed-by: Fangzhi Zuo <jerry.zuo@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
This DC patchset brings improvements in multiple areas. In summary, we have:
- Increased KUnit coverage
- Reverts a patch causing regression on DCN42
- Fixes for cursor
- Imporvements for ABM
Reviewed-by: George Zhang <george.zhang@amd.com>
Signed-off-by: Taimur Hassan <Syed.Hassan@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[WHY]
element_size_to_bytes_per_pixel() handled only element sizes 0 to 4 and
returned 0 for every other value, and an unexpected element size caused
a divide by zero.
[HOW]
Derive the size as 1 << element_size, which matches the existing
encoding, and fall back to 1 for values of 8 or more to keep the result
nonzero.
Reviewed-by: Austin Zheng <austin.zheng@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
This reverts commit 555641725f21 ("drm/amd/display: Request DMUB HW cursor offload").
[WHY]
This commit causes the following failures on DCN42:
igt@kms_plane_cursor@viewport
igt@kms_plane_cursor@overlay
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Qi <Haoming.Qi@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why&How]
The cursor programming path caches the last-written cursor state in SW
and, for the enable bit, skips the register write when the cache already
matches the request. Under cursor offload the SW cache is updated even
though the actual register write is deferred to firmware. When an
offloaded cursor update is aborted (e.g. an overlay geometry change
reprograms the cursor directly in the same frame), the queued firmware
write is dropped but the cache is left stale, so the following direct
programming pass skips the enable write and the cursor disappears over
the overlay plane.
Add a refresh_cursor_state() hubp/dpp hook that reads the cursor
position and attribute registers back from hardware into the SW cache,
and call it from dcn35_abort_cursor_offload_update() so the cache is
truthful after an abort. This keeps the enable-change optimization
intact for all ASICs and is race-robust regardless of whether firmware
applied the queued write before the abort. Only dcn60 (the sole offload
user) wires up the hook.
Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why&How]
There can be more than 2 DCFCLK values for DCN6, so the assert is not
applicable.
Remove it.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
dcn10_build_cursor_pos_update_params
This commit contains renames, adding const to pipe_ctx in
build_cursor_pos_update_params and also removing unnecessary checks
before dc_dmub_srv_is_cursor_offload_enabled().
[Why]
It improves code clarity.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Piotr Maziarz <piotr.maziarz@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
dcn401_build_cursor_position() clamped x_hotspot to 0xFF before
publishing *pos_out, but that field is also stored into hubp->curs_pos
and forwarded to dpp1_set_cursor_position(), which derives
src_x_offset from it. The ODM/MPC slice adjustments earlier in the same
function deliberately grow x_hotspot past 255 to keep the cursor
visible across a slice boundary, so the clamp corrupted DPP source
addressing in exactly the case it was meant to handle.
[How]
Clamp in hubp401_cursor_set_position() instead, where the 8-bit
CURSOR_HOT_SPOT_X field lives. A local is used for the register write
only; pos->x_hotspot passes through unmodified.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Piotr Maziarz <piotr.maziarz@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
UT mirror headers still declared the old prototypes and hwss_gtest still
called the callback directly. Separately, the divide by
param.pixel_clk_khz used to sit after hubp401_cursor_set_position()'s
"curs_attr.address == 0" early return; hoisting the math into the
builder moved it ahead of that guard, so it now traps on streams with
no timing programmed.
[How]
Guard the dst_x_offset block with "if (param.pixel_clk_khz)"
Route the test through hwss_program_cursor_position().
Update dc_hwss_ut.h / dc_hwseq_ut.h and the BLS docs to the new signatures.
Move the pos/param declarations to the top of the loop body.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Piotr Maziarz <piotr.maziarz@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Flatten the executor params for the initial set of HWSS block sequence ops
for SET_CURSOR_POSITION to separate SW and HW logic
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Rafal Ostrowski <rafal.ostrowski@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
pipe_ctx is not allowed to reach a block sequence step. Several operations
still carried it into the executor, which then walked the logical state at
execute time (pipe_ctx->stream->timing, pipe_ctx->plane_state and similar)
instead of acting on data resolved during build.
[How]
Give each operation a flat parameter list holding only the hardware objects
and values the final hwss call needs, and change the matching hooks to take
it. The function pointer, the params struct, the executor, the builder
helpers and all call sites are updated together.
Operations covered:
- OPTC_PROGRAM_MANUAL_TRIGGER - takes tg
- HUBP_PROGRAM_TRIPLEBUFFER - program_triplebuffer takes (hubp, bool)
- DPP_SETUP_DPP, DPP_PROGRAM_BIAS_AND_SCALE - take (dpp, plane_state)
- DSC_ENABLE_WITH_OPP - takes (dsc, opp_inst)
- OPP_PROGRAM_BIT_DEPTH_REDUCTION - opp
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Rafal Ostrowski <rafal.ostrowski@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[WHY&HOW]
Adds diagnostic logs for recent DMU traces.
Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Reviewed-by: Aric Cyr <aric.cyr@amd.com>
Signed-off-by: Dillon Varone <Dillon.Varone@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
Inbox1 was previously being used for cursor and DIG locking with
DMUB doing the relevant programming on behalf of driver.
With the Inbox0 lock the driver is expected to do the programming
itself and the Inbox0 lock acts as nothing more than a SW mutex.
There are still a few gaps in how Inbox1 was being leveraged in the
cursor path that need to be addressed for Inbox0.
[How]
Add a `should_use_dmub_inbox0_lock_for_link` helper to help
check if we support the Inbox0 locking infrastructure and we should
be locking it if the specific link in question requires it.
Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Zhikai Zhai <zhikai.zhai@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct clock_source and struct dio. The DCE110
and DCE112 clock sources derive it from their clock_source_id; the
DCN10 DIO is a singleton and uses 0.
Also move the (void)encoding cast in dcn401_program_pix_clk() below the
local declarations, where it belongs.
No functional change; the inst field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct link_encoder, struct stream_encoder and
struct hpo_frl_stream_encoder, and populate it in the DCN10 through
DCN60 construct paths.
No functional change; the field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Description]
Re-enable P-State message to SMU now that support has been added on
SMU side
Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Alvin Lee <Alvin.Lee2@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Description]
When sending the indicate P-Stat message to PMFW, the wait_resp flag
should be set based on fw_based_mclk_switching in order to maintain
the logic from previous product.
Reviewed-by: Wenjing Liu <wenjing.liu@amd.com>
Signed-off-by: Alvin Lee <Alvin.Lee2@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct pg_cntl and set it to 0 in
pg_cntl35_create() and pg_cntl42_create(), since PG cntl is a singleton.
No functional change; the field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
HWS_APPLY_UPDATE_FLAGS_FOR_PHANTOM and HWS_UPDATE_PHANTOM_VP_POSITION
program no hardware. They only mutate driver state: pipe update flags,
the phantom plane viewport and the phantom scaling params.
A block sequence is built first and executed later, so both ops ran only
after the builder had already inspected pipe_ctx->update_flags to decide
which steps to emit. The phantom pipe was therefore built from stale
flags.
[How]
Call apply_update_flags_for_phantom() and update_phantom_vp_position()
directly from hwss_build_post_unlock_full_sequence(), before
program_pipe_sequence(), matching the legacy
dcn401_post_unlock_program_front_end() ordering.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Piotr Maziarz <piotr.maziarz@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
Hwss executors shouldn't be coupled to complex dc structs such as
pipe_ctx. The DPP_PROGRAM_UPSP executor took a whole pipe_ctx only to
reach two fields.
[How]
Change program_upsp_params to hold the dpp and the dscl_prog_data that
dpp_program_upsp() actually needs, and extract them in the generic HWSS
layer. Add hwss_add_dpp_program_upsp() so the build site uses a helper
with the standard bounds check instead of assigning the step inline.
Update the BLS specs.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Piotr Maziarz <piotr.maziarz@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct dmub_replay and set it to 0 in
dmub_replay_construct(), since the replay object is a pool singleton.
Also move the (void)panel_inst cast in dmub_replay_set_coasting_vtotal()
below the local declarations, where it belongs.
No functional change; the inst field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct dmub_psr and set it to 0 in
dmub_psr_construct(), since the PSR object is a pool singleton.
No functional change; the field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct abm and set it to 0 in dce_abm_construct()
and dmub_abm_construct(), since ABM is a singleton.
No functional change; the field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
CalculateFlipSchedule has slightly different calculations depending on
if it was called by mode support or mode programming.
Mode support calculates the lower bound of the required bandwidth for
immediate flips. This takes into the account of the register limits to
ensure it is still programmable.
Mode programming uses the required bandwidth for immediate flips on a
per-plane basis to calculate the number of dst_y lines needed to achieve
this bandwidth.
Usually the immediate flip bandwidth per plane determined by mode
programming is above the lower bound calculated by mode support.
However, mode programming can fail if that isn't the case since a
bandwidth lower than the lower bound will end up exceeding the register
limit used by hardware. e.g. The proportion of a single plane's flip
bandwidth w.r.t to the total bandwidth available for all planes can
result in one of the planes having a lower BW than the lower bound
calculated in mode support
[How]
Update function to always use the use_lb_flip_bw path. Add calculations
for dst_y_per_vm/row_flip lines. Consolidate calculation of lb_flip_bw
by having it consider register limits: Instead of doing two steps of: 1.
Taking the max of the bandwidths without considering register limits. 2.
Taking the max of itself with the bandwidths with register limits
accounted for. Do it in one step as it is easy to miss either step 1 or
2.
Also add some local variables for intermediate values and remove no
longer referenced inputs to improve readability.
Also increase DCN401 max_flip_time_lines by 2 lines to ensure previously
supported configs remain supported as worst-case bandwidth can increase
compared to before.
Fixes: 1e719006b623 ("drm/amd/display: Unify CalculateFlipSchedule Logic")
Fixes: 285a363f8c63 ("Revert "drm/amd/display: Unify CalculateFlipSchedule Logic"")
Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Austin Zheng <Austin.Zheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member to struct clk_mgr and set it to 0 in the DCN401,
DCN42, DCN42B and DCN60 construct paths, since clk_mgr is a singleton
on all of them.
No functional change; the field has no readers yet.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Description]
Double buffering needs to be enabled for DSCCLK updates otherwise
a frame of underflow could occur when updating the DSC pipe config
or DSC clock.
Reviewed-by: Ilya Bakoulin <ilya.bakoulin@amd.com>
Signed-off-by: Alvin Lee <Alvin.Lee2@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add an inst member and initialize it to 0 in every hubbub*_construct()
from DCN10 through DCN60.
No functional change.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add the AMD Open-Source SPDX line to dc files.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why&How]
DCN6 has 4 pipes, so we should be able to use 3 overlay + 1 primary plane.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why & How]
1. The periodic ABM full-frame update for VESA PR is now owned
by the DMU timer.
2. In PR_update_state, compute whether the current scenario allows
the periodic keep-alive (not CTS, not front-buffer rendering,
not force-full-frame, content actively updating) and forward it
via the abm_periodic_ffu_allowed runtime flag only when it changes.
Reviewed-by: Robin Chen <robin.chen@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Leon Huang <Leon.Huang1@amd.com>
Signed-off-by: George Zhang <george.zhang@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
dcn30_apply_idle_power_optimizations() derives the MALL frame cache
hysteresis timer with
tmr_delay = (uint32_t)(div_u64(..., denom) - 64LL);
div_u64() returns a u64, so when the quotient is smaller than 64 the
subtraction wraps instead of going negative and tmr_delay ends up huge.
The loop that follows tries to squeeze it into the 6 bit register field
by doubling denom, but that only makes the quotient smaller, so tmr_delay
can never converge. tmr_scale is bumped past 3 and the function gives up
with
/* Delay exceeds range of hysteresis timer */
ASSERT(false);
even though the requested delay is too *short* to encode, not too long.
With mall_additional_timer_percent left at its default of 0, the quotient
drops below 64 once the refresh rate used for the calculation goes above
~243 Hz. Every DCN 3.0 display above that loses MALL static screen
entirely and splats a WARN once per boot. Reproduced on Navi 23
(RX 6600) driving 1920x1080, resetting /sys/kernel/debug/clear_warn_once
between modes:
refresh MALL ASSERT
144 Hz enabled no
240 Hz enabled no
280 Hz skipped yes
360 Hz skipped yes
Commit 3bb68cec4db8 ("drm/amd/display: Add Overflow check to skip MALL")
already covered the other end of the range, where a large stutter period
makes the delay too long to encode. Cover the short end by clamping to
0, which selects the shortest hysteresis the register can express,
65.28us * 64 = ~4.18ms. That is marginally longer than what the formula
asks for at these refresh rates, and erring long is the safe direction:
it only delays MALL entry, it can never enter early.
The numerator does not change between iterations, only denom does, so
compute it once and keep both call sites inside 100 columns.
The genuinely out of range case at very low refresh rates still reaches
the ASSERT, which is where it belongs.
Fixes: 52f2e83e2fe5 ("drm/amdgpu/display: add MALL support (v2)")
Signed-off-by: Francis Marlou Pacaro <pacaro.francis.marlou.n@gmail.com>
Reviewed-by: Leo Li <sunpeng.li@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
dcn42_prepare_bandwidth
dc->clk_mgr is already dereferenced inside dcn401_prepare_bandwidth(),
so the dc->clk_mgr NULL check that follows is redundant and misleading.
Drop it.
Fixes: 53845307d52e ("drm/amd/display: Register DCN as a PMFW DF C-state client on DCN42")
Reported-by: Dan Carpenter <error27@gmail.com>
Cc: Roman Li <roman.li@amd.com>
Cc: Alex Hung <alex.hung@amd.com>
Cc: Tom Chung <chiahsuan.chung@amd.com>
Cc: Ivan Lipski <ivan.lipski@amd.com>
Cc: Harry Wentland <harry.wentland@amd.com>
Cc: Aurabindo Pillai <aurabindo.pillai@amd.com>
Cc: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Cc: Wayne Lin <wayne.lin@amd.com>
Cc: Dan Wheeler <daniel.wheeler@amd.com>
Cc: Leo Chen <leo.chen@amd.com>
Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
dc->clk_mgr is checked for NULL earlier in dcn50_init_hw() and
dcn60_init_hw(), but dcn50_initialize_min_clocks() and
dcn401_initialize_min_clocks() are called without any guard,
causing Smatch to report potential NULL dereferences.
Guard both call sites with the same pattern used throughout
both functions:
if (dc->clk_mgr && dc->clk_mgr->funcs)
Also fix dcn50_initialize_min_clocks() which calls
get_dispclk_from_dentist without checking the function pointer,
unlike the dcn401 equivalent which guards that call.
Fix kernel-doc in dcn60_hwseq.c by adding missing parameter descriptions
for @probe in dcn60_update_probe_status() and @type in
is_probe_measurement_type_for_hubbub().
Fixes: 7f7d7ea1fa51 ("drm/amd/display: Add new sources for DCN6")
Reported-by: Dan Carpenter <error27@gmail.com>
Cc: Aurabindo Pillai <aurabindo.pillai@amd.com>
Cc: Ivan Lipski <ivan.lipski@amd.com>
Cc: Dan Wheeler <daniel.wheeler@amd.com>
Cc: Roman Li <roman.li@amd.com>
Cc: Alex Hung <alex.hung@amd.com>
Cc: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Fix the spelling in comments. No functional change.
Signed-off-by: Ruslan Vagner <rusya92266@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
This version brings along the following updates:
- Expand amdgpu_dm KUnit coverage across connector init, sink stream
creation, FreeSync caps, forced atomic commit, plane modifiers, cursor
updates, panic flush, MST connector creation, AUX transfers, link
bandwidth readback and MST DSC configuration.
- Refactor RMCM into a separate module.
- Decouple cursor offload hwss executors from pipe context.
- Decouple HUBP_UPDATE_PLANE_ADDR from pipe_ctx.
- Rename lock_and_validation_needed to needs_dc_state_realloc.
- Drop the dead update_type param from update_planes_and_stream_adapter.
- Attach only plane updates and stream updates that actually changed.
- Unify fast update classification paths.
- Request DMUB HW cursor offload.
- Program DCC as part of address update.
- Update LLS and UPSP programming paths.
- Add lock-free memory pool.
- Add instance field to struct mpc and struct dccg.
- Cleanup DMUB command submission interfaces.
- Add inbox0 HW lock helpers for DCN35.
- Atomize IRQ register read/modify/write ops.
- Add Replay cumulative residency query.
- Add urgent assertion counter probe.
- Add debug option to force optional UCLK support.
- Force DSC to 8bpp for SST and MST DP tunneling over USB4.
- Honor forced RGB pixel encoding.
- Add option for certain panels to disable FEC.
- Cap DML2.1 vmin ODM combine at 2:1 for eDP.
- Add is_odm_enabled callback to skip init_odm on active ODM pipes.
- Remove MALL capabilities from DCN42B.
- Skip MALL calculations when there is no MALL.
- Enable power gating on dcn42b.
- Remove SDPIF_PORT_CONTROL programming for DCN31/35/42.
- Enable back alt-ch.
- Flush ISM work before releasing the stream.
- Fix HDMI FRL audio enable.
- Fix peak bandwidth measurement sequence.
- Bound the DSC power gating loop by num_dsc.
- Cast DP DTO pixel clock math to avoid overflow and narrowing.
- Use unsigned types for FRL cap check params and HPO read_state.
- Return success status from check_mode_supported.
- Add SPDX license identifier to dcn30_dpp_cm.c.
Acked-by: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Taimur Hassan <Syed.Hassan@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
Commit c3fb1fb9e65f ("drm/amd/display: Fix warnings") aligned the
signedness of a large number of DC values, but two hunks of the original
change were not carried over:
- struct frl_cap_chk_params_fixed31_32 still declares audio_packet_type,
h_active and h_blank as int. All three only ever hold non-negative
values: h_active and h_blank are HDMI timing quantities counted in
pixels, and audio_packet_type is only compared against the positive
HDMI audio packet type constants 0x02, 0x07, 0x08, 0x09 and 0x0e in
frl_capacity_computations_common().
- hpo_enc3_read_state() still declares pixel_encoding, color_depth and
odm_combine as int and passes their addresses to REG_GET_2()/REG_GET(),
whose generic_reg_get*() backends take uint32_t *. The mismatch is only
hidden by the explicit (uint32_t *) cast inside the REG_GET macros. The
DCN401 equivalent, hpo_enc401_read_state(), already uses uint32_t.
[How]
Widen the three struct fields to unsigned int/uint32_t and the three
locals to uint32_t so the types match how the values are produced and
consumed. No computed value changes.
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
The dc_fast_update intermediate struct created code duplication and
complexity with multiple classification paths (populate_fast_updates,
fast_nonaddr_updates_exist, full_update_required). This refactoring
simplifies the update classification system by consolidating to a
single path while maintaining compatibility.
[How]
Remove entire dc_fast_update struct and associated helper functions:
- populate_fast_updates
- fast_nonaddr_updates_exist
- full_update_required
Refactor check_update_surfaces_for_stream as the
single classification path with explicit handling for func_shaper,
lut3d_func, cursor_csc_color_matrix_change, and
scaler_sharpener_update. Add a reserved bitfield to the
stream_update_flags union for completeness guards. Extract
dc_check_address_only_update and dc_check_update_surfaces_for_stream
as public.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Rafal Ostrowski <rafal.ostrowski@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
Add dcn35_dmub_hw_control_lock() and dcn35_dmub_hw_control_lock_fast(),
a SW lock mechanism to prevent racing between driver and FW.
The helpers are not wired into dcn35_funcs/dcn351_funcs yet, so there is
no functional change: DCN3.5/3.5.1/3.6 keep using the existing inbox1
lock path. Enabling them will come in a follow-up patch.
Reviewed-by: Ray Wu <ray.wu@amd.com>
Signed-off-by: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
Commit 0421fc6ab3a8 ("drm/amd/display: fix __udivdi3 link error") switched
get_dp_dto_frequency_100hz() and dcn401_get_dp_dto_frequency_100hz() to
div_u64() but dropped the original explicit casts, leaving two issues:
- In get_dp_dto_frequency_100hz(), "clock_hz * dp_dto_ref_khz * 10" is
computed in 32-bit unsigned arithmetic (both operands are unsigned int)
before being stored into the u64 temp, so the product can overflow
before it is widened (flagged by Coverity OVERFLOW_BEFORE_WIDEN).
- div_u64() returns a u64 that is assigned directly to the unsigned int
*pixel_clk_100hz, an implicit narrowing conversion that trips
-Wconversion (possible loss of data) on stricter builds.
[How]
Cast clock_hz to unsigned long long so the multiplication is performed in
64-bit, and make the u64 -> unsigned int narrowing explicit with an
(unsigned int) cast on the div_u64() results in both functions. The
computed values are unchanged.
Fixes: 0421fc6ab3a8 ("drm/amd/display: fix __udivdi3 link error")
Reviewed-by: Wayne Lin <Wayne.Lin@amd.com>
Signed-off-by: James Lin <PingLei.Lin@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why/How]
The APG_CLOCK_EN bit can stay at 0 in some cases when going between
DPMS off and DPMS on.
Add APG_CLOCK_EN to audio_mute_control functions to make sure it gets
programmed.
Reviewed-by: Alvin Lee <alvin.lee2@amd.com>
Signed-off-by: Ilya Bakoulin <Ilya.Bakoulin@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why & How]
DCN42 has 0 bytes allocated for MALL, but this results in a 0 <= 0 comparison in
CalculateMALLUseForStaticScreen(), which blindly returns true.
This causes DML to calculate with is_using_mall_for_ss[n] = true.
Example:
0 + 0 <= 0 -> true
|
v
is_using_mall_for_ss[1] = true
|
v
use_one_row_for_frame[1] = true
PTE_BUFFER_MODE[1] = true
|
v
dpte_row_height : 128 -> 808
dpte_row_width_ub : 196,608 -> 1,204,224
|
+--> DST_Y_PER_PTE_ROW_NOM_L = 808
|
+--> N_groups = ceil(1,204,224 / 65,536 / 2) = 10
|
v
REFCYC_PER_PTE_GROUP_NOM_L = 8416 (0x20E0)
Currently, DML does not plumb out the programming for PTE_BUFFER_MODE and FORCE_ONE_ROW_FOR_FRAME,
but even adding that programming does not enable the mode to work with these values.
The simplest fix is to just block any opportunistic MALL calculations.
Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Ovidiu Bunea <ovidiu.bunea@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Description]
dml21_top_utm_check_mode_supported must return the success status from
the mode support result otherwise the caller will believe that mode
support always returns true. This can cause issues where mode support
returns true, where mode programming returns false.
Reviewed-by: Wenjing Liu <wenjing.liu@amd.com>
Signed-off-by: Alvin Lee <Alvin.Lee2@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
DCN42B has no MALL, but MALL capabilities are non-zero.
[How]
Zero out MALL capabilities in the dcn42b bounding box.
Reviewed-by: Matthew Stewart <matthew.stewart2@amd.com>
Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Gabe Teeger <gabe.teeger@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
DCN42B has no MALL, but the resource file advertised MALL and CAB
sizes. SubVP was also enabled even though its phantom surfaces
are allocated in MALL.
[How]
Drop the MALL and SubVP caps, the dml2_options svp_pstate and
mall_cfg setup, and the add_phantom_pipes and
calculate_mall_ways_from_bytes callbacks. Disable SubVP.
Reviewed-by: Matthew Stewart <matthew.stewart2@amd.com>
Signed-off-by: Gabe Teeger <gabe.teeger@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why&How]
dcn30_dpp_cm.c carries the full MIT license text but no SPDX tag. Add
the SPDX-License-Identifier line so the file matches the rest of the
tree.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|
|
[Why]
struct dccg has no instance identifier, unlike other DC hardware
blocks. Debug and tracing helpers need one to identify the block.
[How]
Add an inst member to struct dccg and initialize it to 0 in every
dccgX_create(), since there is one DCCG instance per ASIC.
Reviewed-by: Jun Lei <jun.lei@amd.com>
Signed-off-by: Tony Cheng <Tony.Cheng@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
|