DualPathOcc: Dual-Resolution BEV Encoder for 3D Occupancy Prediction
本文提出DualPathOcc框架,通过高分辨率特征聚合和双路径BEV编码器解决多视图图像的3D占用预测问题,优化模型以提高预测精度。
本文提出DualPathOcc框架,通过高分辨率特征聚合和双路径BEV编码器解决多视图图像的3D占用预测问题,优化模型以提高预测精度。
本文提出CPR-IE指标,通过表征经济性、预测质量和资源负担来比较智能系统在部署限制下的表现,解决了仅依靠预测准确性进行比较的问题。
研究通过虚拟现实实验探讨了自动驾驶卡车与行人在多车道道路上的互动风险,并测试了三种干预措施,发现投影式外部人机界面最有效。
研究构建了CMPM基准,用于评估视觉-语言模型对中国多面板梗图的理解能力,通过两个任务测试模型对结构类型、顺序敏感度及梗图解释生成的性能。
Traditional quantum image generation methods are constrained by pixel-position encoding, causing quantum resource requirements to scale with image resolution and inducing probabilistic competition among jointly decoded pixels, which hinders precise control. This work proposes a coordinate-conditioned implicit generation paradigm: the image is modeled as an implicit function driven by spatial coordinates and latent variables, where a classical embedding network generates parameters for a variational quantum circuit. Pixel intensities are obtained independently at each coordinate by evaluating the circuit and measuring the expectation values of dedicated color qubits. This approach decouples resolution from the number of address qubits, circumvents inter-pixel probability normalization constraints, and incorporates structural inductive bias. Experiments demonstrate that the proposed framework achieves superior visual and quantitative performance on two benchmark datasets compared to FRQI, PQWGAN, and classical baselines, using significantly fewer qubits.
本文提出DualPathOcc框架,通过高分辨率特征聚合和双路径BEV编码器解决多视图图像的3D占用预测问题,优化模型以提高预测精度。
本文提出CPR-IE指标,通过表征经济性、预测质量和资源负担来比较智能系统在部署限制下的表现,解决了仅依靠预测准确性进行比较的问题。
研究通过虚拟现实实验探讨了自动驾驶卡车与行人在多车道道路上的互动风险,并测试了三种干预措施,发现投影式外部人机界面最有效。
研究构建了CMPM基准,用于评估视觉-语言模型对中国多面板梗图的理解能力,通过两个任务测试模型对结构类型、顺序敏感度及梗图解释生成的性能。
Traditional quantum image generation methods are constrained by pixel-position encoding, causing quantum resource requirements to scale with image resolution and inducing probabilistic competition among jointly decoded pixels, which hinders precise control. This work proposes a coordinate-conditioned implicit generation paradigm: the image is modeled as an implicit function driven by spatial coordinates and latent variables, where a classical embedding network generates parameters for a variational quantum circuit. Pixel intensities are obtained independently at each coordinate by evaluating the circuit and measuring the expectation values of dedicated color qubits. This approach decouples resolution from the number of address qubits, circumvents inter-pixel probability normalization constraints, and incorporates structural inductive bias. Experiments demonstrate that the proposed framework achieves superior visual and quantitative performance on two benchmark datasets compared to FRQI, PQWGAN, and classical baselines, using significantly fewer qubits.